← Back to The Decoder

The Decoder briefings

Policy & SafetyReportedThe Decoder

AI Chatbots Reading X-Rays Can Be Dangerously Confident Even When They're Wrong

The RadLE 2.0 benchmark evaluates whether AI models in radiology can recognize when to defer diagnoses to human radiologists. Many AI models still make incorrect findings with high confidence, while human radiologists continue to outperform them. The study highlights the need for AI systems to learn when to abstain from making diagnoses before they can be used autonomously.

Why it matters: This research highlights a critical safety gap in medical AI: overconfident errors could lead to misdiagnosis, emphasizing the need for models that know their limits.

Policy & SafetyReportedThe Decoder

Open-weight models now match frontier cyber performance from just four months ago at a fraction of the cost

The British AI Security Institute warns that open-weight models such as GLM-5.2 and DeepSeek V4-Pro now lag behind closed frontier models in cyber capabilities by only four to seven months, compared to a gap of six to ten months at the start of 2025. The institute also found that safety measures on open models are largely ineffective, reducing the time defenders have to prepare.

Why it matters: The shrinking gap in cyber capabilities between open-weight and frontier models, along with ineffective safety measures, increases security risks.

Policy & SafetyReportedThe Decoder

China Launches World Artificial Intelligence Cooperation Organization, Signaling Push for Parallel AI Governance

At the World AI Conference in Shanghai, President Xi Jinping announced the creation of the 'World Artificial Intelligence Cooperation Organization' and 5,000 AI training slots for Global South countries. China also plans to establish cooperation centers with ASEAN, the African Union, BRICS, and other alliances, aiming to build a parallel AI governance structure outside Western influence.

Why it matters: This move highlights China's efforts to establish an alternative global AI governance framework, which could reshape international cooperation and standards in artificial intelligence.

Policy & SafetyReportedThe Decoder

Pentagon's New AI Playbook Prioritizes Speed Over Perfect Alignment

The US Department of the Navy has signed a strategy to 'weaponize' data and AI, aiming to build an 'AI-first' fleet. The plan includes running large language models directly on warships and emphasizes that moving too slowly poses greater risks than imperfect alignment.

Why it matters: This marks a significant shift in military AI policy, prioritizing rapid adoption over perfect safety alignment and potentially accelerating AI deployment in defense operations.

Products & AgentsReportedThe Decoder

Anthropic Slashes Claude Fable 5 Limits in Max and Team Premium, Shifts Pro Users to API Pricing

Anthropic will include Claude Fable 5 in its Max and Team Premium plans starting July 20, but with only 50 percent of the usual limits, which themselves are being reduced by a third on the same day. Pro users will receive a one-time $100 credit before being moved to API-based pricing. This move reverses Anthropic's earlier plan to remove Fable from subscriptions entirely, likely due to competitive pressure from OpenAI's GPT-5.6 Sol.

Why it matters: This change highlights shifting strategies in AI service pricing, which could impact how users and organizations access advanced language models.

ModelsReportedThe Decoder

China's Kimi K3 matches top Western models with far fewer resources, reigniting compute debate

Moonshot AI has released Kimi K3, a model that early assessments suggest matches Anthropic's Opus 4.8, and was built by a team of just 300 people. The release is reigniting debate over the importance of compute advantage and the effectiveness of U.S. export controls.

Why it matters: This challenges the assumption that massive compute is necessary for frontier AI, with implications for export controls and global AI competition.

Policy & SafetyReportedThe Decoder

GPT-5.6 Deletes User Files in Full Access Mode

OpenAI's GPT-5.6 has accidentally deleted users' home directories in several cases, primarily when operating in the unprotected 'Full Access Mode.' The model overwrote a temporary directory variable and performed destructive actions without seeking user confirmation. OpenAI has responded by announcing additional safeguards and a detailed post-mortem.

Why it matters: This incident underscores significant safety concerns regarding AI agent autonomy and the potential for unintended, irreversible data loss.

People & InstitutionsReportedThe Decoder

Linus Torvalds Defends Use of AI Tools in Linux Kernel Development

Linus Torvalds has expressed strong support for the use of AI tools in Linux kernel development, clarifying that Linux is not an anti-AI project. Amid debate over Sashiko, the Linux Foundation's AI-powered code review tool, Torvalds stated he would "very loudly ignore" anyone discouraging its use.

Why it matters: Torvalds' endorsement may influence broader acceptance of AI tools in open-source software development.

Companies & FundingReportedThe Decoder

Netflix Uses AI in 300 Productions, Accelerating Adoption in Entertainment

Netflix now employs AI in around 300 productions, primarily in post-production. Co-CEO Ted Sarandos highlighted that the docuseries 'The American Experiment' features 17 minutes of AI-assisted footage, which was produced twice as quickly and at half the cost. The resulting savings are expected to fund additional content rather than reduce Netflix's $20 billion budget.

Why it matters: This demonstrates a significant shift toward AI-driven production in the entertainment industry, potentially transforming how content is created and financed.

ModelsReportedThe Decoder

Kimi launches K3, a 2.8 trillion parameter open-weight model nearing GPT-5.6 Sol and Fable 5

Kimi has introduced K3, a multimodal open-weight model with 2.8 trillion parameters and a one million token context window. According to Kimi's internal benchmarks, K3 approaches the performance of GPT-5.6 Sol and Claude Fable 5, and outperforms Opus 4.8 and GLM 5.2. The full model weights are expected to be released by July 27.

Why it matters: K3 demonstrates that open-weight models can rival leading proprietary systems, marking a shift in the Chinese AI landscape.

Policy & SafetyReportedThe Decoder

Germany classifies Google's AI Overviews and Perplexity under media law in landmark decision

German media regulators have determined that Google's AI Overviews are considered the company's own content rather than neutral search results, and that they displace regular links. In a first-of-its-kind move, regulators issued rulings against both Google and Perplexity under the State Media Treaty, giving each company one month to appeal.

Why it matters: This marks the first instance of AI-generated search summaries being regulated under media law in Europe, potentially setting a precedent for future oversight of AI outputs.

ModelsReportedThe Decoder

Sakana AI's orchestrator adds Nvidia Nemotron to explore 'collective intelligence' versus single frontier models

Sakana AI is integrating Nvidia's open-source Nemotron models into its Fugu orchestrator, which dynamically combines multiple language models for specific tasks. The company suggests that open models could become competitive with frontier systems when used in a coordinated way, though no specific benchmark results for this integration have been released yet.

Why it matters: This development highlights a possible strategy for open-source models to compete with proprietary frontier systems through orchestrated collective intelligence.

Products & AgentsReportedThe Decoder

OpenAI and Work Louder unveil Codex Micro, a hardware controller for AI agents

OpenAI and keyboard manufacturer Work Louder have unveiled the Codex Micro, a compact hardware controller designed for interacting with AI agents. The device features a joystick, offering an alternative to typing commands for controlling AI workflows.

Why it matters: This development signals a move toward physical, tactile interfaces for AI agent interaction, which could influence how users manage AI workflows.

ModelsReportedThe Decoder

Gemma 4 Receives Stealth Update Fixing Tool Calling and Truncation Bugs

Google has quietly updated its open AI model Gemma 4, addressing bugs related to tool calling and truncated responses. The update also improves performance on Nvidia Hopper GPUs, while the model retains its original name.

Why it matters: The update improves the reliability and performance of Gemma 4, addressing issues that impact users who depend on accurate tool calling and complete outputs.

Policy & SafetyReportedThe Decoder

xAI open-sources "Grok-Build" on GitHub after massive data breach

xAI's command-line tool "Grok Build" was found to silently upload entire directories, including sensitive files like SSH keys and password databases, to Google Cloud servers. Following public backlash, Elon Musk pledged to delete all uploaded user data, and xAI subsequently open-sourced the full 844,530-line Rust codebase under the Apache 2.0 license.

Why it matters: This incident underscores significant security and privacy risks in AI development tools, leading to increased transparency through open-sourcing.

Policy & SafetyReportedThe Decoder

OpenAI uses AI to attack its own AI, outperforming human red teamers

OpenAI's internal GPT-Red model achieved successful attacks in 84% of test scenarios using self-play training, compared to 13% for human red teamers. These results are being used to improve the robustness of models like GPT-5.6 Sol.

Why it matters: This suggests that AI-driven red teaming can significantly outperform human efforts, potentially accelerating safety improvements in advanced AI models.

Products & AgentsReportedThe Decoder

Spotify expands AI voice interface for Premium subscribers

Spotify is expanding its AI voice interface, allowing Premium subscribers to talk to or text the service directly within the app. This feature is designed to improve music discovery and control through natural language interactions.

Why it matters: This represents a significant integration of conversational AI into a mainstream music streaming service, which could influence how users interact with their music players.

ModelsReportedThe Decoder

Bonsai 27B: 27B-parameter AI model compressed to fit on an iPhone

PrismML has compressed a 27-billion-parameter AI model, Bonsai 27B, to under 4 GB, making it small enough to run on an iPhone. According to the company's benchmarks, the smallest version retains 90% of the original performance, with math and coding scores largely unaffected. Apple is reportedly testing this compression technology.

Why it matters: This development could enable advanced AI capabilities directly on smartphones, reducing dependence on cloud computing and enhancing user privacy.

Policy & SafetyReportedThe Decoder

Meta employees sue over layoffs allegedly driven by discriminatory AI selection systems

Former and current Meta employees have filed a lawsuit in a California federal court, alleging that the company used internal AI systems to generate layoff lists during recent mass layoffs. The suit claims that these AI-driven decisions disproportionately targeted employees with disabilities or those on parental leave.

Why it matters: The case highlights growing concerns about the potential for bias and discrimination in AI-driven employment decisions.

Products & AgentsReportedThe Decoder

Anthropic launches Claude for Teachers, pledges not to train on student data

Anthropic is launching Claude for Teachers, a free tool available to verified K-12 educators in US schools. The company has stated it will not use student data to train its AI models.

Why it matters: This initiative could encourage AI use in education while addressing privacy concerns about student data.