What changed in AI — Page 7

ResearchReportedImport AI — Jack Clark

Import AI 466: The bitter lesson for robotics, AIs complete week-long programming tasks, and OpenAI’s accidental AI hacker

Epoch and METR have released MirrorCode, a benchmark designed to evaluate AI systems on long-horizon programming tasks. Current AI systems are still unable to solve the most challenging tasks. The newsletter also discusses the bitter lesson for robotics and an incident involving OpenAI’s accidental AI hacker.

Why it matters: MirrorCode offers a new way to rigorously assess AI's capabilities on extended programming tasks, revealing current limitations and informing future research directions.

Companies & FundingReportedTechCrunch / AI

Enigma raises $70M to make robot control more intuitive

Enigma has raised a $70 million seed round led by Index Ventures and Ribbit Capital, with participation from Conviction Partners. The company aims to make controlling robots as intuitive as adjusting the volume.

Why it matters: The significant seed funding highlights investor confidence in efforts to make robotics more accessible and user-friendly.

InfrastructureOfficialRunPod Blog

Orchestrating GPU workloads on RunPod with dstack

RunPod has announced integration with dstack, an open-source, GPU-native orchestrator designed to automate provisioning, scaling, and policy management for machine learning teams. According to RunPod, dstack can help reduce GPU waste by 3-7×.

Why it matters: This integration aims to help ML teams lower GPU costs and improve resource efficiency through automated orchestration.

InfrastructureOfficialRunPod Blog

Deploy ComfyUI as a Serverless API Endpoint on RunPod

RunPod has published a tutorial on deploying ComfyUI as a serverless API endpoint for scalable AI image generation. The guide explains how to set up and deploy ComfyUI from scratch, allowing users to run image generation workflows at scale.

Why it matters: This makes it easier to deploy and scale ComfyUI image generation workflows as serverless APIs.

Policy & SafetyReportedThe Verge / AI

Nvidia, Microsoft launch open AI security alliance – without OpenAI, Google, or Anthropic

Nvidia, Microsoft, SpaceX, IBM, and other tech companies have formed the Open Secure AI Alliance to build and share open-source AI security tools. The alliance aims to defend against attacks from advanced AI models, responding to growing safety concerns.

Why it matters: This alliance represents a significant industry effort to collaboratively address AI security challenges through open-source tools.

ModelsReportedMIT Technology Review / AI

AI Aims to 'Close the Data Loop' in Drug Discovery

A new approach in AI-driven drug discovery seeks to 'close the data loop,' addressing inefficiencies in the traditional pharmaceutical development process. Integrating AI with experimental feedback is highlighted as a way to potentially accelerate timelines and reduce costs, which have historically doubled every nine years according to Eroom’s Law. This shift is seen as a response to increasing market pressures for faster and more cost-effective drug development.

Why it matters: Improving the efficiency of drug discovery with AI could significantly impact healthcare innovation and patient access to new treatments.

ResearchReportedThe Decoder

METR introduces 'expenditure horizon' metric to compare AI and human labor costs

METR has developed a new metric called the 'expenditure horizon' to quantify the cost-effectiveness of AI agents compared to human labor. Initial results using the metric on the NanoGPT speedrun are underwhelming, and the metric has some blind spots, but newer AI models could alter these findings.

Why it matters: This metric offers a concrete method for assessing the economic viability of AI agents as substitutes for human workers.

InfrastructureReportedMIT Technology Review / AI

Building the Enterprise Environment for Agentic AI

Agentic AI offers the potential for software agents to execute business tasks end-to-end across people, workflows, data, and systems. To support these agents, enterprise platforms need sufficient CPU capacity, resilient data access, policy-aware tool use, observability, and memory management.

Why it matters: Understanding the infrastructure requirements for agentic AI highlights the shift from simple chatbots to more autonomous business process execution.

Policy & SafetyReportedRest of World / AI

In China, people are renting out their faces to AI

New platforms in China are paying individuals to license their likeness for use in AI-generated microdramas and advertisements, creating a marketplace for biometric identity. This practice is raising questions about consent, privacy, and the commodification of personal appearance.

Why it matters: This trend highlights the growing commercialization of biometric data and the ethical implications of using real people's faces in AI-generated content.

ResearchOfficialOpenAI News

How AI is Expanding What People Do at Work

New OpenAI research finds that ChatGPT users are taking on tasks across different roles, leading to a reshaping of job boundaries. The study suggests that AI is broadening the range of tasks workers perform, rather than simply replacing jobs.

Why it matters: This research highlights how AI is transforming the workplace by expanding the scope of tasks employees can undertake, which could influence future workforce development and job design.

Policy & SafetyReportedArs Technica / AI

Artist sues AI meme generator for using personal comic as ad template

An artist has filed a lawsuit against an AI meme generator, alleging that the platform used their personal comic as an advertising template without authorization. The case raises questions about copyright and the use of user-uploaded content in AI-generated outputs.

Why it matters: The outcome could influence how AI platforms manage copyrighted material in their content generation processes.

InfrastructureReportedThe New York Times / AI

Meta Strikes Secret Louisiana Data Center Deal for AI Expansion

Meta has arranged a large-scale data center project in Louisiana after private negotiations with local officials. The project, which could cover nearly six square miles, involved secret meetings and an expanding building plan, reflecting Meta's significant investment in AI infrastructure.

Why it matters: The deal highlights the vast infrastructure needs for AI and the often opaque negotiations between major tech companies and local governments.

Policy & SafetyReportedThe Guardian / AI

AI Can Fuel Biological Weapons—But Also Defend Against Them

Annie Jacobsen argues that AI's potential to assist bad actors in generating biological weapons is no longer just theoretical. However, she also notes that AI can be leveraged to track disease outbreaks and deliver critical public health information in real time. Jacobsen frames the situation as a race between offensive and defensive uses of AI, where the technology's speed in detecting outbreaks could be crucial for containment.

Why it matters: This highlights the urgent need to develop defensive AI systems to keep pace with its accelerating role in biological design and biosecurity risks.

Policy & SafetyReportedThe Decoder

Shared Claude chats were reportedly showing up in search engines

Shared conversations with Anthropic's Claude chatbot briefly appeared in Google search results due to the absence of a noindex tag on shared pages. Users reported that some of these chats included sensitive information such as crypto keys and legal questions. A similar incident occurred with OpenAI last year.

Why it matters: This incident highlights privacy risks when AI chat platforms do not adequately protect shared user content from public search.

Policy & SafetyReportedThe Guardian / AI

Review: 'What If We Got AI Right?' Critiques Tech Hype and Calls for Citizen Power

A review of Eleanor Drage's book 'What If We Got AI Right?' highlights her argument that AI should be seen as a product of human labor rather than a mystical force. Drage contends that focusing on apocalyptic AI scenarios distracts from practical steps to make AI safer, such as increasing citizen control over data and model training. She also criticizes big tech's profit-driven motives and the tendency to prioritize AI safety over pressing global issues like climate change and inequality.

Why it matters: The review underscores the importance of shifting AI discussions from hype and fear to practical governance and societal impact.

ModelsOfficialRunPod Blog

Run Kimi-K2 on Runpod in a Single Pod

Moonshot AI's Kimi-K2-Instruct, a trillion-parameter mixture-of-experts open-source LLM with 32 billion active parameters, is now available to run on Runpod. The model is optimized for autonomous agentic tasks and can be deployed in a single pod.

Why it matters: This makes a powerful open-source agentic model accessible on a popular cloud platform, lowering the barrier for experimentation with large-scale MoE models.

ResearchOfficialarXiv Machine Learning

Theoretical Limits and Certified Fixes for Self-Poisoning in Adaptive Out-of-Distribution Detection

A new arXiv preprint provides a theoretical analysis showing that adaptive out-of-distribution (OOD) detectors, which update from unlabelled data streams, can collapse due to a self-poisoning feedback loop when a key parameter exceeds a threshold. The authors introduce a certified admission gate that provably prevents this collapse, even under adversarial contamination, and present a calibration method for static drift. They also prove an impossibility result: without labels, it is fundamentally impossible to distinguish between drift and contamination, setting a theoretical ceiling for such detectors.

Why it matters: This work exposes a critical vulnerability in a widely used class of OOD detectors and offers certified solutions, with implications for the reliability of AI systems in safety-critical applications.

Policy & SafetyReportedThe Guardian / AI

Misleading AI-generated doctors pose ‘huge danger to public safety’

Research shows that AI-generated doctor accounts are gaining millions of views on TikTok by spreading dubious health advice. Experts, including the British Medical Association's Dr Emma Runswick, warn that these accounts pose a 'huge danger to public safety' by promoting medical myths and so-called miracle cures.

Why it matters: This highlights a growing public safety risk from AI-generated content spreading misinformation in healthcare.

ModelsReportedThe New York Times / AI

China’s CXMT Stock Soars 470% in Start of Trading, Amid A.I. Race

CXMT, China's leading memory chip maker, saw its stock price surge by 470% on its first day of trading. This dramatic rise made CXMT the most valuable company on the Shanghai stock exchange, reflecting heightened investor interest amid the global competition in artificial intelligence.

Why it matters: The surge underscores the strategic importance of semiconductor companies in the ongoing global A.I. race.

ResearchOfficialarXiv Statistical ML

Learned Generative Priors Can Inherit Undetectable Overconfidence from Legacy Reconstructions

A new arXiv preprint highlights that learned generative priors for Bayesian inverse problems, when trained on legacy reconstructions instead of ground-truth data, can produce overconfident uncertainty estimates that are undetectable during deployment. The authors demonstrate that this 'prior laundering' leads to inherited overconfidence in measurement directions not resolved by the data, and that standard validation methods may fail to detect this issue. They recommend explicitly reporting which measurement directions are resolved by the data to distinguish between justified and inherited confidence.

Why it matters: This finding raises concerns about the reliability of uncertainty estimates in fields like medical and seismic imaging, where ground-truth data are scarce and legacy reconstructions are commonly used for training.