What changed in AI — Page 15

Companies & FundingReportedAI Business

UK Robot Maker Humanoid Valued at $1.35 Billion

UK-based robotics company Humanoid has reached a valuation of $1.35 billion after a recent funding round. This development highlights the company's emergence as a notable European player in a robotics sector largely led by Chinese and U.S. firms.

Why it matters: This signals the rise of a significant European competitor in the global robotics market, which has been dominated by Chinese and American companies.

Policy & SafetyReportedThe Verge / AI

Lawmakers prepare bill requiring AI ‘kill switch’

Lawmakers are preparing to introduce an 'AI Kill Switch Act' that would require AI companies to shut down or throttle their systems on orders from the Department of Homeland Security. Representatives Ted Lieu (D-CA) and Nathaniel Moran (R-TX) are expected to introduce the legislation on Thursday.

Why it matters: The bill could give the government significant authority over AI systems, raising questions about safety, oversight, and regulatory power.

ResearchOfficialIBM Research

IBM Research Collaborates on Unified AI Benchmarking Platform

IBM Research is collaborating with a global team to make AI benchmarking results easier to compare, replicate, and reuse. The initiative seeks to address fragmentation in AI evaluation by developing a unified platform for sharing and accessing benchmarking data.

Why it matters: Standardized benchmarking is important for tracking AI progress and enabling fair comparisons across different models.

Products & AgentsOfficialElevenLabs Blog

ElevenLabs Adds References Feature for Music v2

ElevenLabs has introduced References, a new feature for its Music v2 model that allows users to upload an existing track to guide the style and feel of AI-generated music. This enhancement enables users to achieve more precise control over the sound of their creations.

Why it matters: References gives creators the ability to closely match the style and mood of reference tracks in their AI-generated music.

ModelsReportedMIT Technology Review / AI

AI Accelerates Design of Next-Generation Medicines

Artificial intelligence is increasingly being used to assist scientists in designing new medicines, especially biologic therapies made from engineered proteins. By leveraging AI, researchers can more efficiently identify promising drug candidates, potentially reducing the time and cost associated with traditional drug development.

Why it matters: AI-driven drug design could accelerate the development of innovative treatments and improve patient outcomes.

ModelsReportedThe Decoder

Google CEO Pichai says Gemini's next leap depends on building 'much larger base models'

Alphabet has raised its 2026 investment forecast to as much as $205 billion, citing demand outpacing spending. Google Cloud reportedly grew 82% in the second quarter. CEO Sundar Pichai stated that Google needs a larger base model for its next AI leap and has initiated an ambitious Gemini 4 training run.

Why it matters: This highlights Google's commitment to scaling up its AI models, which could drive significant advancements in AI capabilities and influence industry competition.

ModelsReportedMarkTechPost / AI

Best Open Speech Recognition (ASR) Models in 2026: WER, Languages, Latency, and License Compared

A roundup compares 16 open-weight ASR models on word error rate, language coverage, streaming latency, and license. Models such as Cohere Transcribe, IBM Granite Speech 4.1, ARK-ASR, and MOSS-Transcribe are separated by less than one WER point on the Hugging Face Open ASR Leaderboard, indicating that rank alone is no longer decisive. The article also explains why published averages cannot be directly subtracted from one another.

Why it matters: Open speech recognition now features multiple competitive models, making model selection more nuanced than in the previous Whisper-dominated landscape.

InfrastructureReportedMarkTechPost / AI

Meet Gigatoken: A Rust BPE Tokenizer that Encodes Text at 24.53 GB/s, up to 989x Faster than HuggingFace Tokenizers

Gigatoken is an MIT-licensed Rust BPE tokenizer that achieves encoding speeds of 24.53 GB/s on a 144-core AMD EPYC 9565 processor. It outperforms HuggingFace tokenizers by up to 989x and tiktoken by 681x, with speed gains attributed to a hand-written SWAR pretokenizer and pretoken caching rather than a faster BPE merge loop.

Why it matters: This speedup could significantly reduce preprocessing time for large-scale language model training and inference.

Products & AgentsReportedMarkTechPost / AI

Anthropic Releases Claude Security Plugin for Claude Code in Beta

Anthropic has released the Claude Security plugin for Claude Code in beta. The plugin performs a multi-agent vulnerability scan of a repository from within an existing Claude Code session, then generates patch files for selected findings for users to review and apply.

Why it matters: This plugin integrates automated security scanning into the developer workflow, potentially reducing vulnerabilities in codebases.

Products & AgentsReportedMarkTechPost / AI

Cursor Releases Cursor Router: A Request-Level Classifier Delivering Frontier Coding Quality at 30–50% Lower Cost

Cursor has made Cursor Router generally available for Teams and Enterprise plans. The system classifies each request by query, context, task complexity, and domain, then routes it to the most suitable model. Cursor reports frontier-quality output at 60% savings in online A/B tests, and 30–50% savings for three early-access enterprise accounts measured against Opus 4.8 rates.

Why it matters: Cursor Router offers a practical way to reduce coding AI costs without sacrificing quality, which could make advanced AI-assisted development more accessible to enterprises.

Open SourceReportedMarkTechPost / AI

Unsloth vs Axolotl vs TRL vs LLaMA-Factory: A Fine-Tuning Framework Comparison on Speed, VRAM, and Multi-GPU

A comparison of four open-source LLM fine-tuning frameworks—Unsloth, Axolotl, TRL, and LLaMA-Factory—shows their distinct engineering priorities. Unsloth focuses on kernel rewrites for speed, Axolotl on parallelism strategies, TRL on trainer APIs, and LLaMA-Factory on broad model coverage. The article explores trade-offs in speed, VRAM usage, and multi-GPU support.

Why it matters: This comparison informs developers about the strengths of each framework, aiding in selecting the best tool for specific fine-tuning requirements.

ModelsReportedMarkTechPost / AI

Cisco Foundation AI Releases Antares: 350M and 1B Open-Weight Models for Vulnerability Localization in Codebases

Cisco Foundation AI has released Antares, a family of small language models designed to identify the locations of known vulnerabilities within codebases. Antares-1B achieves a 0.209 File F1 score on the new Vulnerability Localization Benchmark, outperforming much larger models such as GLM-5.2 (753B parameters) and Gemini 3 Pro. Running a full 500-task benchmark sweep takes about 13 minutes on a single H100 GPU for under a dollar, compared to $141 for GPT-5.5.

Why it matters: Antares shows that small, specialized models can surpass much larger general-purpose models in targeted security tasks, offering a more cost-effective solution for vulnerability localization.

ModelsReportedMarkTechPost / AI

Poolside Releases Laguna S 2.1, an Open-Weight Agentic Coding Model Punching Above Its Weight Class on SWE-Bench Multilingual

Poolside has released Laguna S 2.1, a 118B open-weight Mixture-of-Experts coding model with 8B active parameters per token and a 1M-token context window. The model matches or outperforms much larger models on agentic coding benchmarks and can run on a single NVIDIA DGX Spark.

Why it matters: Laguna S 2.1 demonstrates that high performance in agentic coding tasks is possible with fewer active parameters, potentially lowering hardware requirements.

Companies & FundingOfficialIBM Research

IBM acquires HRL Laboratories

IBM has acquired HRL Laboratories, a research lab recognized for its pioneering work in technologies such as the laser and silicon spin qubits. The acquisition is intended to enhance IBM's research capabilities in quantum computing and other advanced fields.

Why it matters: The deal is expected to strengthen IBM's position in quantum computing and advanced research by leveraging HRL's expertise and legacy.

ResearchReportedThe New York Times / AI

Could A.I. Do Your Job? We Put Agents to the Test.

The New York Times conducted an experiment deploying AI agents as office workers. The agents were able to complete some assigned tasks, but struggled or failed with others.

Why it matters: This experiment offers insight into the current strengths and limitations of AI agents in real-world office environments.

Open SourceOfficialHugging Face Blog

Bringing Nunchaku 4-bit Diffusion Inference to Diffusers

Hugging Face has integrated Nunchaku, a 4-bit quantization method for diffusion models, into the Diffusers library. This allows for more efficient inference of models such as FLUX.1-dev and SD3.5, reducing memory usage and potentially speeding up generation. The integration is available as an open-source tool.

Why it matters: This development lowers hardware requirements for high-quality image generation by enabling efficient 4-bit quantized inference.

Policy & SafetyOfficialarXiv Software Engineering

Widespread License Laundering Found in AI Dataset and Model Supply Chains

A large-scale analysis of over 230,000 dataset-to-model-to-application chains on Hugging Face and GitHub finds that 62.3% of these chains include at least one artifact lacking a declared license. The study also reports that licenses with legal obligations rarely persist through the supply chain, with less than 7% end-to-end survival, while permissive licenses are retained in 95.1% of cases.

Why it matters: This highlights a significant challenge for legal compliance and rights enforcement in the AI ecosystem, raising concerns for developers, rights holders, and platform operators.

Companies & FundingReportedTechCrunch / AI

ServiceNow invests $40 million in Indian banking software firm BusinessNext to boost global AI banking push

ServiceNow has invested $40 million in BusinessNext, an Indian banking software specialist, at a $700 million valuation. The partnership aims to help expand BusinessNext's AI-powered banking software globally, strengthening ServiceNow's presence in financial services.

Why it matters: The deal highlights ServiceNow's commitment to expanding its AI-driven financial services offerings and global reach.

ResearchOfficialarXiv Software Engineering

LLMs Struggle to Detect Faulty Code When Documentation Remains Intact, Study Finds

A new arXiv preprint introduces TRACE, a method for evaluating how large language models (LLMs) allocate trust across conflicting software artifacts such as code, documentation, and tests. Testing seven LLMs on Java method bundles with injected faults, the study finds that models are much better at detecting errors in documentation than in code, and often fail to deprioritize faulty implementations when documentation appears correct. The results suggest LLMs may over-rely on documentation and miss subtle code errors, with confidence scores offering little help in distinguishing correct from incorrect judgments.

Why it matters: This highlights a significant limitation in current LLM-based coding assistants, raising concerns for their use in safety- or correctness-critical software development.

ResearchOfficialarXiv Software Engineering

New CEO-Bench Benchmark Shows AI Agents Struggle with Long-Term Startup Management

A new arXiv preprint introduces CEO-Bench, a benchmark designed to test language model agents on managing a simulated startup over 500 days, requiring long-term planning, adaptation, and multi-task coordination. The study finds that only a few advanced models—Claude Fable 5, GPT-5.6 Sol, and Claude Opus 4.8—end with more than the initial $1M balance, but all perform worse than a simple rule-based baseline. This highlights a significant gap in current AI agents' ability to handle complex, sustained decision-making tasks.

Why it matters: The results suggest that even leading AI models are not yet capable of reliably managing complex, long-term real-world tasks, underscoring a key limitation for practical deployment.