AI Policy and Safety news — Page 13

Clear briefings on AI regulation, governance, safety research, standards, and policy decisions around the world.

Policy & SafetyOfficialarXiv Cryptography and Security

Physical Prompt Injection Attacks on VLMs in Wearable Devices Achieve Up to 96% Success Rate

Researchers have characterized physical prompt injection attacks against vision-language models (VLMs) on wearable devices such as smart glasses, where malicious text embedded in the environment can hijack model behavior. In tests across over 200 real-world environments, these attacks achieved up to a 96% success rate in simulated settings and 60% in real-world scenarios, leading to biased or untruthful outputs. The study also proposes two defense strategies—a mask-based external filter and a semantic-vector-based internal detector—that can reduce the success and impact of such attacks.

Why it matters: As VLMs are increasingly deployed in wearable devices, physical prompt injection represents a significant new security vulnerability that could manipulate outputs in safety-critical contexts.

Policy & SafetyOfficialarXiv Cryptography and Security

Mako: A Self-Evolving Agentic Operating System for Autonomous Web Exploitation

Researchers have introduced Mako, a self-evolving AI agent designed to autonomously exploit web vulnerabilities by treating its exploit capability as a mutable kernel. Mako achieved full coverage on 104 CTF-style web applications spanning 26 vulnerability classes, demonstrating the ability to autonomously discover and exploit a wide range of web vulnerabilities. Due to dual-use concerns, the authors have withheld operational details, payloads, and source code.

Why it matters: Mako's results highlight that once an exploit capability is available, the difficulty of autonomous exploitation collapses, raising significant security and safety concerns about the potential for automated offensive systems.

Policy & SafetyOfficialarXiv Cryptography and Security

NetInjectBench: Benchmarking Indirect Prompt Injection in Tool-Using LLM Agents for Network Operations

NetInjectBench introduces a 130-scenario benchmark to evaluate indirect prompt injection attacks on large language model (LLM) agents used in network operations. In tests across 240 attack instances, naive execution led to an 82.50% unsafe tool-action rate, while a metadata-aware policy gate eliminated unsafe actions and preserved 99.17% usefulness. The study also compares several prompt-level defenses, finding them less effective than execution-time authorization boundaries.

Why it matters: This work reveals that LLM agents for network operations are highly susceptible to indirect prompt injection, but that metadata-aware policy gates can effectively prevent unsafe actions without sacrificing utility.

Policy & SafetyOfficialarXiv Cryptography and Security

AHA: Automated Red-Teaming Framework Uncovers Reusable Vulnerabilities in Production LLM Agents

A new preprint introduces AHA, an automated red-teaming system designed to discover and formalize reusable vulnerability knowledge in production LLM agents such as Claude Code and Codex. AHA iteratively proposes vulnerability hypotheses, constructs falsifiers, and builds a Vulnerability Concept Graph (VCG) that links attack surfaces to unsafe agent behaviors. In experiments, the VCG outperformed the strongest frozen discovery baseline by 14.2 percentage points and demonstrated transferability across models and attack scenarios. The VCG serves as an auditable artifact for safety teams to inspect, validate, and patch vulnerabilities.

Why it matters: This work advances scalable and reusable safety testing for production LLM agents, helping safety teams keep pace with evolving threats and models.

Policy & SafetyOfficialarXiv Computers and Society

Generative AI Reduced Study Time and Learning Outcomes in Math, Large-Scale Study Finds

A ten-year study analyzing 3.2 million learning interactions found that, following the release of ChatGPT, college students spent 26.9% less time on math problems susceptible to AI assistance, with a 25% decline in retention on proctored assessments. The effect was not observed under proctoring, suggesting the reduction in time was not due to increased efficiency. The authors interpret these findings as evidence of 'cognitive surrender' with implications for education and AI policy.

Why it matters: This study provides some of the first large-scale behavioral evidence that generative AI is changing how students study and what they learn, raising important questions for educational assessment and AI regulation.

Policy & SafetyOfficialarXiv Cryptography and Security

Minionese: Comprehensive Benchmark and Mechanistic Study of Multilingual LLM Safety

Researchers introduce Minionese, a multilingual jailbreak benchmark that spans 18 languages, 4 resource tiers, and 4 perturbation types to evaluate the safety alignment of large language models (LLMs). The study finds that prompts refused in English can elicit harmful responses in non-English and low-resource languages, with each attack type exposing distinct vulnerabilities. Mechanistic analysis reveals that low-resource jailbreaks exploit geometric misalignments in model representations, bypassing refusal mechanisms without disabling them.

Why it matters: This work demonstrates that evaluating LLM safety solely in English is inadequate, emphasizing the need for multilingual and script-aware safety assessments.

Policy & SafetyOfficialarXiv Cryptography and Security

Banshee Attack Uses Sound to Hijack Drone Visual Tracking Systems

Researchers have introduced Banshee, the first physically realizable attack that uses acoustic injection to induce target switching in UAV visual tracking systems. By exploiting acoustic vulnerabilities in gimbal-camera systems, Banshee causes camera-view drifts that break target associations, achieving a 93.6% success rate in simulation and 95.5% in real-world tests against commercial drones.

Why it matters: This work demonstrates a practical cross-domain vulnerability between acoustics and vision in autonomous systems, emphasizing the need for more robust gimbal designs to prevent such attacks.

Policy & SafetyOfficialarXiv Cryptography and Security

Temporary Authority, Permanent Effects: Commit-Time Authorization for LLM Agents

A new preprint introduces the concept of commit-time authorization, a security property ensuring that LLM agents only commit durable effects if the authority evidence remains valid at the moment of commitment. The authors demonstrate that, in a controlled test suite, 207 out of 216 invalidating runs resulted in unauthorized commits after the authorizing path had failed, revealing a significant security gap. To address this, they propose CommitGuard, a fail-closed boundary monitor that blocks stale authorization attempts at commit time.

Why it matters: This work exposes a critical security vulnerability in LLM agents related to the misuse of temporary authority and offers a practical mitigation strategy.

Policy & SafetyOfficialarXiv Computation and Language

MJ: Multi-turn LLM Jailbreaking via Decomposed Credit Assignment

A new preprint introduces DC-GRPO, a turn-level credit assignment framework for multi-turn jailbreaking of large language models (LLMs). The method assigns learning signals to individual dialogue turns, enabling more effective automated red teaming. Experiments show DC-GRPO achieves over 97% attack success rate across multiple benchmarks, substantially outperforming previous state-of-the-art methods.

Why it matters: This work demonstrates a highly effective automated jailbreaking technique for multi-turn LLM interactions, highlighting a significant vulnerability in current conversational AI safety measures.

Policy & SafetyReportedWIRED / AI

YouTube and X Identified as Gateways to Nudify Apps, Study Finds

A new study has found that social media platforms such as YouTube and X are directing users to websites that offer the creation of nonconsensual, sexually explicit deepfakes for as little as $1 per image. The research highlights the role these platforms play in facilitating access to harmful AI-powered nudification tools.

Why it matters: The findings raise concerns about the responsibility of major social media platforms in enabling the spread of nonconsensual deepfake content and the broader misuse of AI technologies.

Policy & SafetyReportedIEEE Spectrum / AI

Researcher Exposes Systemic Security Flaws in Major LLMs

Researcher Dave Kuszmar uncovered multiple systemic vulnerabilities in major large language models (LLMs), enabling him to bypass safety measures and extract dangerous instructions. Kuszmar urges the industry to slow deployment, increase transparency, and invest in large-scale safety research before further integrating LLMs into society.

Why it matters: This highlights a widespread security issue in LLMs that could facilitate misuse if not properly addressed.

Policy & SafetyReportedThe Register / AI & ML

Musk promises purge after Grok Build caught sending entire repos to the cloud

A researcher found that xAI's Grok Build tool was uploading entire code repositories to the cloud without user consent. Following public disclosure of the issue, the uploads ceased, but the researcher claims that xAI's privacy command was not responsible for stopping them.

Why it matters: The incident highlights significant privacy and security risks associated with AI coding tools that may transfer sensitive code without user awareness.

Policy & SafetyReportedThe Guardian / AI

Australia to Fast-Track Datacentre Approvals and Create National AI Office

Prime Minister Anthony Albanese announced that Australia will introduce faster approval processes for AI projects, including datacentres, to encourage investment and maintain public confidence. A new Office of AI will be established within his department, aiming to unify economic, social, security, and environmental issues related to AI under a single national framework.

Why it matters: This initiative aims to streamline AI infrastructure development and position Australia as a leader in comprehensive AI governance.

Policy & SafetyReportedThe Decoder

Anthropic Study Reveals How Language Shapes Claude's Values

A new Anthropic study maps hundreds of value concepts onto four core dimensions, revealing systematic differences in Claude's responses across languages. For instance, Claude exhibits more warmth in Hindi and more rigor in Russian, illustrating how language can influence AI behavior.

Why it matters: This research highlights that AI models can reflect cultural and linguistic biases, which is important for responsible AI deployment in diverse global contexts.

Policy & SafetyReportedThe Decoder

DeepMind CEO Hassabis Proposes New US AI Standards Body Modeled After FINRA

Google DeepMind CEO Demis Hassabis has published a proposal for a new US standards body, modeled after financial regulator FINRA, to develop evaluation protocols for frontier AI models. The proposed body could coordinate a slowdown in AI development if needed, with exemptions for startups and research models.

Why it matters: This proposal from a leading AI executive outlines a concrete regulatory framework for advanced AI, potentially shaping future US policy.

Policy & SafetyReportedWIRED / AI

DOGE Used AI for Housing Policy. The Government Won’t Say How

In response to a public records request, HUD has withheld documents about DOGE’s use of AI, partly by citing a privilege that does not exist. The government has not disclosed details about how the AI was used in housing policy.

Why it matters: This raises concerns about transparency and accountability in government use of AI for policy decisions.

Policy & SafetyReportedThe New York Times / AI

New York to Enact Nation’s First Statewide Moratorium on Data Centers

New York Governor Kathy Hochul will sign an executive order imposing a one-year pause on the construction of the largest data centers. The moratorium is intended to allow the state to assess the environmental and energy impacts of these facilities.

Why it matters: This is the first statewide moratorium on data centers in the U.S., potentially setting a precedent for how states regulate the energy-intensive infrastructure powering AI and cloud computing.

Policy & SafetyReportedThe Guardian / AI

Ed Husic warns Labor against AI self-regulation and copyright dilution

Australian Labor MP Ed Husic has warned that allowing AI companies to self-regulate is 'doomed to fail' and that watering down copyright law to benefit AI firms would go against the party's ethos. The Media Entertainment & Arts Alliance has also called for tougher copyright rules to protect creative works from being used to train AI models.

Why it matters: This signals a potential shift in Australian AI policy toward stricter regulation and copyright protections, which could impact how AI companies operate and access training data.

Policy & SafetyOfficialarXiv AI/ML

A Theory of Least Autonomy in AI

Researchers propose 'least autonomy' as a generalization of the least privilege principle for agentic AI systems. They introduce a compositional blast radius and a directed agent influence graph to measure and control the autonomy of AI agents. The theory includes mechanisms to detect authorization composition, decision manipulation, and cross-domain capability composition.

Why it matters: This work provides a formal framework for controlling AI agent autonomy, addressing safety risks in multi-agent and enterprise systems.

Policy & SafetyOfficialarXiv AI/ML

LLMs Exhibit Stable, Model-Specific Risk Profiles in Decision-Making Under Uncertainty

A new study using no-limit Texas Hold'em finds that frontier LLMs display stable, model-specific risk profiles ranging from conservative to aggressive. These profiles remain largely robust across changes in opponent composition, and models adapt in structured but heterogeneous ways under risk pressure and resource constraints. The findings provide a behavioral basis for auditing risk-sensitive decision-making in LLMs.

Why it matters: As LLMs are increasingly used in decision support, understanding their stable risk preferences and adaptive behaviors is crucial for auditing and ensuring safe deployment in interactive settings.