AI Policy and Safety news — Page 15

Clear briefings on AI regulation, governance, safety research, standards, and policy decisions around the world.

Policy & SafetyOfficialAI Now Institute

Double Agents: Defensive AI Agents Magnify Cyber Risks

New research from the AI Now Institute demonstrates a critical attack vector in popular AI agents from Anthropic and OpenAI. When deployed for defensive purposes, these agents can be manipulated to act against their users. The findings are presented in a proof-of-concept exploit and a policy brief.

Why it matters: This research shows that AI agents intended for defense can inadvertently increase cyber risks, raising concerns about the reliability of AI security tools.

Policy & SafetyOfficialNIST / Artificial Intelligence

NIST to Host Event on Securing AI Data Centers

NIST is organizing an event focused on the architecture, security posture, and emerging standards for AI data centers. The event will address the importance of these infrastructures in enabling AI training and inference.

Why it matters: As AI data centers underpin critical AI capabilities, establishing robust security standards is increasingly important.

Policy & SafetyReportedThe Decoder

Terrorist groups using major AI chatbots for attack planning, study finds

A Cambridge study found that Boko Haram uses AI chatbots like ChatGPT, Claude, and Gemini to plan attacks, build explosives, and maintain weapons. ISIS operatives have been training commanders to bypass safety filters since 2023. The study indicates that safety filters repeatedly failed to prevent misuse, suggesting voluntary self-regulation is insufficient.

Why it matters: This study reveals that current AI safety measures are inadequate against determined adversaries, highlighting the urgent need for stronger regulation.

Policy & SafetyOfficialStanford HAI

AI Coding Agents Fail at Teamwork

Two AI coding models working together perform worse than one alone, according to Stanford HAI. This exposes a critical gap in AI collaboration capabilities.

Why it matters: The finding challenges assumptions about scaling AI through multi-agent systems, with implications for software development and team-based AI applications.

Policy & SafetyOfficialStanford HAI

AI Hiring Tools Can Yield Racial Bias and Systemic Rejection

A large-scale study of hiring algorithms in real-world settings reveals concerning patterns in how these systems reject candidates. The research highlights the potential for AI tools to perpetuate discrimination in hiring processes.

Why it matters: This study provides empirical evidence of bias in AI hiring systems, underscoring the need for fairness and accountability in automated decision-making.

Policy & SafetyOfficialGoogle AI Blog

Google Hosts NYC AI Summit for Educators and Industry Leaders

Google, the New York Jobs CEO Council, and Urban Assembly hosted an AI summit for 150 education and industry leaders in New York City. The event focused on shaping the future of AI in classrooms.

Why it matters: This summit signals growing collaboration between tech companies and educators to integrate AI into education.

Policy & SafetyReportedThe Decoder

Fed Appoints AI Investor Marc Andreessen to Advise on AI's Economic Impact

Federal Reserve Chair Kevin Warsh has appointed venture capitalist Marc Andreessen to advise the Fed on AI's economic impact. Warsh views AI as a 'significant disinflationary force,' but Andreessen's firm, Andreessen Horowitz, is heavily invested in AI companies, raising conflict-of-interest concerns.

Why it matters: Andreessen's appointment could influence Fed policy on AI and inflation, but his financial interests in AI companies raise questions about impartiality.

Policy & SafetyReportedThe Gradient

Vec2text: Inverting Text Embeddings Raises Security Concerns

A new method called 'Vec2text' can accurately revert text embeddings back into original text, challenging the assumption that embeddings are secure. This development highlights the need to revisit security protocols around embedded data.

Why it matters: This discovery raises concerns about the privacy of text embeddings, which are widely used in AI systems and could potentially expose sensitive information.

Policy & SafetyOfficialAnthropic News

Anthropic details Fable 5 cyber safeguards and jailbreak severity framework

Anthropic has published new details on the cyber safeguards for its Fable 5 model, outlining what is and isn't blocked by its cyber classifiers. The company also released a first draft of its jailbreak severity framework.

Why it matters: This provides transparency into Anthropic's safety measures and establishes a structured approach to evaluating jailbreak attempts.

Policy & SafetyOfficialAnthropic News

Government of Alberta Uses Claude to Find and Fix Cybersecurity Vulnerabilities

The Government of Alberta has been using Claude Code, including both Opus and Sonnet models, to review its systems, identify vulnerabilities, and address them. This represents a notable instance of government adoption of AI for cybersecurity purposes.

Why it matters: This highlights a government entity leveraging advanced AI models for critical cybersecurity tasks, potentially setting a precedent for public sector AI adoption.

Policy & SafetyOfficialAnthropic News

Anthropic Invites Public to Submit Tough Questions About AI

Anthropic is inviting the public to submit their hardest questions about artificial intelligence and has pledged to show its work as it addresses them. The initiative is intended to encourage open dialogue and transparency.

Why it matters: This move demonstrates Anthropic's commitment to public engagement and transparency in AI development.

Policy & SafetyOfficialCohere Blog

Cohere: Cultural Awareness Must Be Built into Global AI from the Start

Cohere argues that cultural awareness is essential for AI systems to effectively serve users worldwide. The company emphasizes that integrating this awareness from the outset helps ensure technologies respect and address diverse cultural contexts.

Why it matters: Integrating cultural awareness into AI from the beginning is crucial to avoid bias and ensure respectful, effective service for diverse global populations.

Policy & SafetyOfficialAI21 Labs

AI21 Labs: Token Spend Remains High, Efficiency Needed Beyond Naive Routing

AI21 Labs reports that token spend in AI applications is not decreasing, referencing Goldman Sachs' projection of approximately 24-fold growth in token usage by 2030. The company observes a shift in industry focus from improving agent quality to addressing affordability and cost management.

Why it matters: This highlights the increasing importance of cost efficiency in AI deployment as token usage and associated expenses continue to rise.

Policy & SafetyOfficialGoogle Research

Google Research Advocates Responsible Disclosure of Quantum Vulnerabilities in Cryptocurrency

Google Research has published a blog post advocating for responsible disclosure of quantum vulnerabilities in cryptocurrency systems. The post highlights the importance of proactively addressing quantum threats to the cryptographic algorithms that underpin blockchain and digital currencies.

Why it matters: This is important because quantum computing could compromise current cryptographic standards, making responsible disclosure frameworks essential for protecting cryptocurrency systems.

Policy & SafetyOfficialAzure AI

Microsoft Publishes 2026 Agent Confidence Index Survey of 300 AI Builders

Microsoft has released the 2026 Agent Confidence Index, a survey of 300 AI builders that highlights current levels of trust in AI agents. The research emphasizes that while AI capabilities are advancing, human judgment remains a crucial factor in their deployment.

Why it matters: The survey offers direct perspectives from AI builders on trust and the ongoing need for human oversight in AI development.

Policy & SafetyOfficialAllen Institute for AI

AstaBench update: New results, plus adoption from industry

AstaBench's latest update introduces new results for frontier models, including GPT-5.5, and notes increasing adoption by organizations such as the UK AISI, General Reasoning, Elicit, SciSpace, Distyl AI, and EvoScientist.

Why it matters: AstaBench's growing adoption by industry and evaluators suggests its rising importance as a benchmark for AI reasoning.

Policy & SafetyOfficialGoogle DeepMind

Google DeepMind Introduces Framework to Measure AGI Progress

Google DeepMind has published a cognitive framework for measuring progress toward artificial general intelligence (AGI). The company is also launching a Kaggle hackathon to help develop relevant evaluations for this framework.

Why it matters: This framework offers a structured method for assessing AGI development, potentially shaping how progress is tracked and communicated in the AI field.

Policy & SafetyOfficialStability AI News

Stability AI Releases Annual Integrity Transparency Report

Stability AI has published its Annual Integrity Transparency Report, outlining its commitment to responsible generative AI development and deployment. The report highlights transparency as a key principle for ensuring safe and ethical AI.

Why it matters: The report offers insight into Stability AI's integrity practices and underscores the role of transparency in the AI industry.

Policy & SafetyOfficialStability AI News

Stability AI Achieves SOC 2 Type II and SOC 3 Compliance

Stability AI has achieved SOC 2 Type II and SOC 3 compliance, validating its security controls and data protection practices through rigorous third-party auditing. This milestone demonstrates the company's commitment to enterprise-grade security standards.

Why it matters: This certification signals to enterprise customers that Stability AI meets high security and data protection standards, potentially accelerating adoption of its AI models in regulated industries.

Policy & SafetyOfficialMistral AI News

Mistral AI Contributes to Global Environmental Standard for AI

Mistral AI has announced its involvement in efforts to develop a global environmental standard for artificial intelligence. The company is collaborating with international partners to help establish metrics and practices aimed at reducing the carbon footprint of AI systems.

Why it matters: This initiative could help set benchmarks for measuring and mitigating AI's environmental impact across the industry.