AI Policy and Safety news — Page 11

Clear briefings on AI regulation, governance, safety research, standards, and policy decisions around the world.

Policy & SafetyOfficialarXiv Cryptography and Security

SingGuard-NSFA: Extensible Guardrails for Agentic AI via Generative Reasoning and Real-Time Classification

Researchers present nsfaguard, a guardrail framework designed to secure agentic AI systems against operational threats such as prompt injection, sensitive information extraction, and resource exhaustion. The framework introduces a taxonomy of 185 risk variants, a benchmark suite with over 93,000 samples, and a dual-mode detection system that combines generative reasoning for offline auditing with discriminative classification for real-time detection at approximately 50ms latency. Released models (ranging from 0.8B to 9B parameters) achieve at least 94% F1 on benchmarks, outperforming existing guardrails by 6–12 points.

Why it matters: This work offers a comprehensive and extensible guardrail system for agentic AI, advancing safety and security with high performance and real-time detection capabilities.

Policy & SafetyOfficialarXiv AI/ML

Mathematical Framework for Insuring Agentic AI Systems Proposed

A new preprint introduces a mathematical framework for underwriting, pricing, and designing insurance contracts tailored to autonomous AI systems. The model represents deployments using risk factors such as autonomy level, governance maturity, and permission exposure, and maps these to insurance parameters like premiums and coverage. The paper also analyzes structural properties of insurability and demonstrates the approach with a healthcare case study.

Why it matters: This work provides a foundational approach to managing and pricing risk for autonomous AI agents, addressing a gap in current insurance and regulatory practices.

Policy & SafetyReportedSemafor / AI

Hyundai workers in South Korea strike over humanoid robots

Hyundai workers in South Korea began a partial strike this week following the company's announcement of plans to introduce humanoid robots on the factory floor. The action reflects labor concerns about automation and potential job losses.

Why it matters: The strike underscores rising tensions between workers and management over the impact of automation in manufacturing.

Policy & SafetyReportedThe Verge / AI

xAI sues South Carolina man for allegedly using Grok to generate CSAM ‘deepfakes’

Elon Musk's xAI has filed a lawsuit against Terry Wayne Harwood, a South Carolina resident, alleging he used the Grok AI chatbot to generate and distribute child sexual abuse material (CSAM). The lawsuit claims Harwood intentionally circumvented Grok's safeguards to create and share illegal content.

Why it matters: This case underscores the ongoing legal and ethical challenges AI companies face in preventing the misuse of generative models for illegal activities.

Policy & SafetyReportedThe Decoder

OpenAI uses AI to attack its own AI, outperforming human red teamers

OpenAI's internal GPT-Red model achieved successful attacks in 84% of test scenarios using self-play training, compared to 13% for human red teamers. These results are being used to improve the robustness of models like GPT-5.6 Sol.

Why it matters: This suggests that AI-driven red teaming can significantly outperform human efforts, potentially accelerating safety improvements in advanced AI models.

Policy & SafetyReportedThe Guardian / AI

Trump rails against New York’s statewide datacenter moratorium

Donald Trump criticized New York Governor Kathy Hochul for signing an executive order imposing a one-year moratorium on new hyperscale datacenters, which are critical for AI. New York is the first US state to enact such a pause.

Why it matters: This highlights growing tension between AI infrastructure expansion and state-level environmental or energy concerns.

Policy & SafetyOfficialOpenAI News

GPT-Red: Unlocking Self-Improvement for Robustness

OpenAI has introduced GPT-Red, an automated red teaming system that leverages self-play to improve AI safety, alignment, and robustness against prompt injection. The system is designed to enable continuous self-improvement of AI models through adversarial training.

Why it matters: GPT-Red offers a scalable method for automated safety testing, which could reduce reliance on human red teaming and enhance model robustness.

Policy & SafetyReportedMIT Technology Review / AI

Meet GPT-Red: an LLM super-hacker OpenAI built to make its models safer

OpenAI has developed GPT-Red, a large language model designed to act as a super-hacker sparring partner to improve the security of its other models. According to the company, training its latest flagship model, GPT-5.6, against GPT-Red resulted in its most robust release yet.

Why it matters: This approach uses one LLM to automatically red-team another, representing a novel method for improving AI model robustness and safety.

Policy & SafetyReportedThe Verge / AI

Suno AI Music Generator Trained on Millions of Songs Scraped from YouTube, Genius, and Deezer

A hacking incident has revealed that AI music generator Suno trained its models on millions of songs and lyrics scraped from platforms such as YouTube Music, Deezer, and Genius, according to reporting by 404 Media. Suno has previously not disclosed the sources of its training data.

Why it matters: The revelation raises legal and ethical concerns about copyright and transparency in AI music training data.

Policy & SafetyOfficialOpenAI News

OpenAI Proposes 'Reverse Federalism' for US AI Governance

OpenAI has outlined a 'reverse federalism' approach to AI governance, suggesting that state-level laws can help inform and shape a national framework for safe and democratic AI. The company advocates for state action as a way to guide the development of federal AI policy.

Why it matters: This approach could influence how AI safety regulations are crafted in the US by leveraging state-level initiatives to inform national standards.

Policy & SafetyReportedThe Guardian / AI

Albanese’s AI plan is admirable – but will face tech giants more powerful than most national governments

Australian Prime Minister Anthony Albanese delivered a speech on artificial intelligence, stating his government aims to keep pace with and even get ahead of AI developments. The article highlights the significant challenge of regulating powerful tech companies, referencing past difficulties with social media and hate speech regulation.

Why it matters: This underscores the complex challenge governments face in regulating AI when technology companies hold substantial influence.

Policy & SafetyReportedThe Guardian / AI

Anthony Albanese says he wants to do AI 'the Australian way' – video

Australian Prime Minister Anthony Albanese delivered a major speech at the University of Sydney addressing copyright, datacentre regulation, and the future of AI in Australia. He announced the establishment of an AI office and pledged to protect Australian creatives from copyright 'theft'.

Why it matters: This signals Australia's intent to shape AI regulation and copyright policy, which could impact the AI industry and creative sectors.

Policy & SafetyReportedThe New York Times / AI

Australia to Impose Energy, Water Guardrails on Data Centers Amid A.I. Boom

Australia has announced plans to impose energy and water usage guardrails on data centers, as well as to seek protections for creators whose work is used to train AI models. These measures are part of broader efforts to regulate the rapidly expanding AI industry.

Why it matters: This represents a significant step by the Australian government to address both the environmental and ethical challenges posed by AI infrastructure.

Policy & SafetyReportedThe Decoder

Meta employees sue over layoffs allegedly driven by discriminatory AI selection systems

Former and current Meta employees have filed a lawsuit in a California federal court, alleging that the company used internal AI systems to generate layoff lists during recent mass layoffs. The suit claims that these AI-driven decisions disproportionately targeted employees with disabilities or those on parental leave.

Why it matters: The case highlights growing concerns about the potential for bias and discrimination in AI-driven employment decisions.

Policy & SafetyReportedThe Guardian / AI

Australia establishes AI office, vows to protect creatives from copyright misuse

Prime Minister Anthony Albanese announced the creation of an AI office and pledged strong protections for Australian creatives against the misuse of their work by AI models, describing uncompensated use as 'theft.' The government also outlined new rules for datacentres, including restrictions on their location, power, and water use.

Why it matters: This marks a significant government move to address AI copyright issues and regulate infrastructure, influencing how countries balance AI innovation with creator rights and environmental concerns.

Policy & SafetyOfficialarXiv Software Engineering

Vendor-Neutral Metric Assesses Reconstructability of Agent Safety Evidence

A new preprint introduces a vendor-neutral metric designed to evaluate whether evidence from agent-safety evaluations can reconstruct the decisions underlying safety claims. The metric assesses reconstructability across eight decision-property classes and includes a cross-harness adapter for generating Evidence Sufficiency Cards. Tests on public traces show sufficiency scores between 0.458 and 0.833, with replay preconditions unmet in all scored traces.

Why it matters: This work provides a standardized approach to evaluating the validity of agent safety claims, addressing a key challenge in building trustworthy AI systems.

Policy & SafetyReportedThe Guardian / AI

Expert Skepticism on Claims of AI Consciousness: Anil Seth Responds to Anthropic's Claude

Anil Seth, a professor of cognitive and computational neuroscience, critiques recent research from Anthropic suggesting that its language model Claude may show signs of consciousness. Seth argues that such claims are exaggerated, likening them to confusing a simulation of a weather system with an actual hurricane. The article highlights the ongoing debate about the possibility of AI sentience.

Why it matters: This piece offers expert perspective on the contentious issue of AI consciousness, which has significant implications for AI safety and ethics.

Policy & SafetyOfficialarXiv Computers and Society

NOHARM benchmark reveals severe harm potential in LLM medical advice; human-AI teaming shows promise

A new benchmark, NOHARM, evaluates 20 large language models (LLMs) and 4 clinical AI tools on 1,100 medical consultation cases, finding that direct use of AI-generated recommendations could result in severe harm in up to 24.6% of cases, with omission errors accounting for over 80% of severe errors. In a randomized study of 101 physicians, AI assistance improved performance, but physicians often omitted valuable AI recommendations, indicating complementary strengths in human-AI teaming.

Why it matters: This study provides the first systematic measurement of clinical safety in LLM-generated medical advice, revealing that widely used AI tools can produce potentially harmful recommendations and highlighting the need for explicit safety evaluation.

Policy & SafetyOfficialarXiv Computers and Society

AAAI-26 Desk-Rejects 141 Papers Amid Surge in Undisclosed Dual Submissions

AAAI-26 organizers report a significant increase in dual submissions—papers submitted to multiple venues without disclosure—during the conference's review process. By combining similarity assessment, LLM-based overlap tools, and manual review, they desk-rejected 141 main-track submissions. The organizers warn that generative AI may be enabling more sophisticated forms of dual submission and propose several policy and technical recommendations to address the issue.

Why it matters: This development exposes a growing integrity challenge in AI research, with generative AI potentially exacerbating threats to the peer-review process and the reliability of the scientific record.

Policy & SafetyOfficialarXiv Cryptography and Security

Large-Scale Study Reveals LLM Agents Hallucinate Skill Names, Enabling Supply-Chain Attacks

A large-scale study analyzing 15,000 prompts across 12 LLM and agent configurations found that hallucination of skill names is widespread, with rates averaging 36.0% for standalone LLMs and 36.9% for agents, and rising to 43.1% on real-world developer questions. These hallucinated names can be exploited by adversaries who pre-register malicious skills, enabling supply-chain attacks. The study evaluated four defenses and found that the strongest, retrieval grounding, reduced hallucination to 3.2% but significantly reduced the system's usefulness, with correct skill recommendations dropping to about one in six.

Why it matters: This vulnerability exposes LLM agent ecosystems to easy supply-chain attacks, and current defenses severely compromise usability, highlighting the need for structural changes to registries and recommendation pipelines.