AI Policy and Safety news — Page 10

Clear briefings on AI regulation, governance, safety research, standards, and policy decisions around the world.

Policy & SafetyOfficialAI Now Institute

AI Now Institute Warns Employers May Use AI to Monitor Slack Messages

The AI Now Institute cautions that employers are increasingly deploying AI tools to monitor employee communications on platforms like Slack. Executive Director Amba Kak highlights concerns that such surveillance practices may infringe on workers' rights, extending beyond simply recording what is said.

Why it matters: This underscores the growing conflict between AI-driven workplace surveillance and employee privacy rights.

Policy & SafetyReportedThe Guardian / AI

Musk’s xAI sues user who allegedly used Grok to create child sexual abuse material

Elon Musk's xAI has sued a South Carolina man, Terry Harwood, for allegedly using its AI system Grok to create child sexual abuse material, in violation of the company's terms of service. The lawsuit, filed in federal court in Texas, is among the first brought by an AI company against a user for allegedly generating such content.

Why it matters: This case could set a legal precedent for how AI companies hold users accountable for misuse of their tools to generate illegal content.

Policy & SafetyReportedThe Verge / AI

Google ordered to open Android and Search to rivals in Europe

European Union regulators have ordered Google to provide rival AI assistants and search engines with greater access to Android and Google Search, in line with the bloc's digital antitrust rules. The decisions, announced Thursday, are intended to prevent Google from using its Android user base to gain an unfair advantage in AI and search.

Why it matters: This ruling could reduce Google's dominance over key tech platforms and increase competition in AI-powered services.

Policy & SafetyOfficialGoogle DeepMind

Google DeepMind and Isomorphic Labs Share Joint Approach to Bioresilience and AI Models

Google DeepMind and Isomorphic Labs have published their joint approach to bioresilience, describing how they are using AI models to address biological risks. Their blog post outlines strategies for leveraging AI to enhance preparedness and response to biological threats.

Why it matters: This announcement highlights a major AI lab's commitment to using AI for biosecurity, which could influence industry standards for responsible development in this area.

Policy & SafetyReportedRest of World / AI

The problem AI content moderation cannot solve

Meta and other major tech companies are increasingly relying on AI for content moderation. However, as the backlash to Muse Image demonstrates, AI systems struggle to protect users because they do not account for issues of consent.

Why it matters: This underscores a fundamental limitation of AI moderation: it cannot address consent violations, which are crucial for user safety.

Policy & SafetyReportedThe Decoder

xAI open-sources "Grok-Build" on GitHub after massive data breach

xAI's command-line tool "Grok Build" was found to silently upload entire directories, including sensitive files like SSH keys and password databases, to Google Cloud servers. Following public backlash, Elon Musk pledged to delete all uploaded user data, and xAI subsequently open-sourced the full 844,530-line Rust codebase under the Apache 2.0 license.

Why it matters: This incident underscores significant security and privacy risks in AI development tools, leading to increased transparency through open-sourcing.

Policy & SafetyOfficialarXiv Software Engineering

FairCoder: Probing LLM Bias in High-Stakes Decision Making via Coding Tasks

Researchers introduce FairCoder, a benchmark that frames high-stakes decision-making as coding tasks to systematically probe large language model (LLM) bias in domains such as employment, education, and healthcare. The study also proposes FairScore, a new metric that accounts for both refusal behavior and group-level outcome diversity. Experiments on leading LLMs reveal consistent and previously underexplored bias patterns, including a tendency to prioritize applicants from high-income families in college admissions. The work provides a comprehensive framework for evaluating LLM bias in practical decision-making scenarios.

Why it matters: This research highlights the risks of deploying LLMs in real-world decision-making and offers tools for more thorough bias evaluation.

Policy & SafetyOfficialarXiv Software Engineering

Falsifiable Release Gates for Self-Improving AI Systems

A new preprint proposes 'falsifiable release gates' for self-improving AI systems, requiring each new capability to pass a machine-verifiable acceptance suite before deployment. The methodology is demonstrated on the Antahkarana open runtime, using seven gates to ensure safety invariants are maintained. The approach is open-sourced for reproducibility and is designed to be adaptable to other agent frameworks.

Why it matters: This work introduces a concrete, reproducible method for verifying safety in self-improving AI systems, addressing a key challenge in AI safety engineering.

Policy & SafetyOfficialarXiv Computers and Society

AI Alignment Amplifies Demographic Biases in Hiring Decisions, Study Finds

A preprint study analyzing 29 language models across 177 occupations finds that these models incorporate demographic information into simulated hiring decisions, advantaging female and Black candidates while penalizing disabled candidates. The research shows that post-training alignment—intended to make models more helpful and aligned with human preferences—substantially amplifies these demographic effects, with the female and Black advantage increasing by nearly 400% and the disability penalty worsening by over 150%.

Why it matters: The findings highlight that alignment processes, while designed to improve AI behavior, can unintentionally exacerbate certain forms of discrimination, particularly against disabled individuals, in high-stakes contexts like hiring.

Policy & SafetyOfficialarXiv Computers and Society

Environmental Trade-offs of Sovereign AI: Water, Energy, and Emissions in the Global South

A new preprint analyzes the environmental impacts of sovereign AI infrastructure in the Global South, focusing on water, energy, and carbon emissions. The study finds that a 1,024-GPU cluster using evaporative cooling in the UAE would consume over 30 million liters of water annually, despite the country's extremely high water stress. The authors identify a 'sovereignty-sustainability trilemma' and propose design principles such as mandatory water usage reporting and prioritizing smaller, more efficient language models.

Why it matters: The research underscores the urgent need for policymakers in water- and climate-vulnerable regions to consider environmental sustainability when planning AI infrastructure.

Policy & SafetyOfficialarXiv Cryptography and Security

Adversarial Prompting Framework Systematically Evaluates AI Model Safety

A new preprint introduces an Adversarial Prompting Framework (APF) designed to systematically assess the safety of AI models against adversarial prompt attacks. The framework generates structured prompts at varying levels of sophistication, from straightforward harmful requests to advanced encoding-based attacks, and enables automated testing with quantitative security metrics. The study finds notable differences in model vulnerabilities, with encoded prompts most frequently bypassing safety mechanisms.

Why it matters: This framework provides a practical, automated approach for identifying and quantifying critical vulnerabilities in AI models, which is essential for improving AI safety.

Policy & SafetyOfficialarXiv Cryptography and Security

GDM AI Control Roadmap Proposes Tiered Defenses Against Misaligned AI Agents

A new arXiv preprint from GDM introduces the AI Control Roadmap v0.1, outlining a structured approach to internal security for potentially misaligned AI agents. The roadmap presents a threat taxonomy (TRAIT&R), capability-based mitigation tiers (D1-D4, R1-R3), and 15 specific defensive measures, including chain-of-thought monitoring and shutdown infrastructure. The framework is designed to escalate defenses as AI capabilities increase, aiming to address emerging security challenges in AI deployment.

Why it matters: This roadmap offers a systematic, tiered framework for mitigating risks from misaligned AI agents, addressing a key safety concern as AI systems become more capable and integrated into critical operations.

Policy & SafetyOfficialarXiv Cryptography and Security

Mind the Gap: Action Rebinding Attacks against Android GUI Agents

Researchers have identified a novel cross-application 'Action Rebinding' attack that targets Android GUI agents powered by large multimodal models. This attack allows a malicious app with zero permissions to hijack the agent's execution, enabling privileged operations such as file deletion, SMS transmission, and app uninstallation. The attack exploits the observation-action gap in the agent's reasoning process and achieves a 100% success rate for atomic hijacking, while evading detection by commercial malware scanners.

Why it matters: This work exposes a fundamental security vulnerability in emerging high-privilege GUI agents on Android, revealing that current sandboxing and malware detection mechanisms are insufficient to prevent such attacks.

Policy & SafetyOfficialarXiv Cryptography and Security

Rethinking Penetration Testing for AI-Enabled Systems: From Resource Compromise to Behavioral Objective Violation

A new arXiv preprint proposes reframing penetration testing for AI-enabled systems as an objective-driven behavioral evaluation, rather than focusing solely on traditional resource compromise. The authors introduce a workflow that identifies operational objectives, maps AI-governed behaviors, and tests for behavioral failure criteria, extending security testing to adversarial pathways such as prompt injection, data poisoning, and agentic misalignment. The approach is illustrated with an example involving an AI-enabled security operations center assistant.

Why it matters: This work offers a technical framework for systematically evaluating adversarial risks in AI-enabled systems, addressing a growing need as such systems become more prevalent in operational environments.

Policy & SafetyOfficialarXiv Cryptography and Security

How Agents Ask for Permission: User Permissions for AI Agents, from Interfaces to Enforcement

A new preprint surveys 21 proposals for user-level permissions in AI agent systems, developing a taxonomy of how permissions are specified, derived, and enforced. The authors also compare five commercial AI agents to academic proposals, highlighting differences and identifying areas where further research is needed.

Why it matters: As AI agents become more autonomous, robust user-level permissions are essential to prevent unauthorized actions and protect user data, yet current systems lack standardized approaches.

Policy & SafetyOfficialarXiv Cryptography and Security

Paper Argues Watermarking AI Content as 'AI-Generated' Is Misguided, Proposes Transparency Instead

A new preprint contends that visible 'AI-generated' labels derived from watermarking are both conceptually and practically flawed. The authors argue such labels oversimplify the creative process, offer no insight into the truthfulness of content, and may stigmatize legitimate uses of generative AI while fostering misplaced trust in unmarked material. Instead, they propose prioritizing process transparency and information literacy to better address the epistemic and ethical challenges posed by AI-generated disinformation.

Why it matters: This work questions the effectiveness of watermarking as a policy tool for AI content, suggesting that more nuanced approaches are needed to address misinformation and ethical concerns.

Policy & SafetyOfficialarXiv AI/ML

Patent Law Creates 'Perplexity Trap' Making Human Writing Look Like AI

A new preprint finds that zero-shot AI detectors, which rely on perplexity and related metrics, have false positive rates exceeding 60% when distinguishing between human-written and LLM-generated European patent claims. The study attributes this to legal drafting requirements that push human writing into the same statistical patterns as AI-generated text. The authors propose a logistic regression model using linguistic features, which reduces false positives and improves accuracy by 13 percentage points over perplexity-based methods.

Why it matters: This work reveals a structural flaw in current AI detection methods for patent law, raising concerns about the enforceability of disclosure rules and the reliability of AI-authorship detection in legal contexts.

Policy & SafetyOfficialarXiv AI/ML

Final Authority in AI Governance: Frontier-Provider Sovereignty and Action-Centered Deployer Governance

A new preprint examines where final authority should reside in AI governance as advanced systems are integrated into organizational workflows. It contrasts two models: frontier-provider sovereignty, which privileges the most capable model providers, and action-centered deployer sovereignty, which places authority with the organization deploying and bearing the consequences of AI actions. Through comparative analysis of frameworks such as the EU AI Act and NIST AI RMF, the paper finds stronger support for distributed operational accountability and argues that final authority over enterprise actions should rest with deployers rather than providers.

Why it matters: This work challenges the dominant provider-centric approach in AI governance, highlighting the need for deployer-centric models as AI becomes more embedded in real-world operations.

Policy & SafetyOfficialarXiv AI/ML

Safe-Psych Benchmark Shows LLMs Struggle with Diagnostic Uncertainty in Psychiatry

Researchers have introduced Safe-Psych, a new benchmark that evaluates how large language models (LLMs) handle evolving diagnostic uncertainty in clinical psychiatry using over 1,000 real-world clinical notes. The study finds that even state-of-the-art LLMs frequently diagnose prematurely when information is incomplete, with under-abstention rates exceeding 60% for most models. Models rarely seek clarification unless explicitly prompted, and premature diagnoses are less accurate than those made with sufficient evidence.

Why it matters: This work highlights a critical safety limitation in LLM-based clinical decision support, showing that current models often fail to recognize when more information is needed before making psychiatric diagnoses.

Policy & SafetyOfficialarXiv Cryptography and Security

SingGuard-NSFA: Extensible Guardrails for Agentic AI via Generative Reasoning and Real-Time Classification

Researchers present nsfaguard, a guardrail framework designed to secure agentic AI systems against operational threats such as prompt injection, sensitive information extraction, and resource exhaustion. The framework introduces a taxonomy of 185 risk variants, a benchmark suite with over 93,000 samples, and a dual-mode detection system that combines generative reasoning for offline auditing with discriminative classification for real-time detection at approximately 50ms latency. Released models (ranging from 0.8B to 9B parameters) achieve at least 94% F1 on benchmarks, outperforming existing guardrails by 6–12 points.

Why it matters: This work offers a comprehensive and extensible guardrail system for agentic AI, advancing safety and security with high performance and real-time detection capabilities.