What changed in AI — Page 13

Policy & SafetyOfficialarXiv Cryptography and Security

Study Finds Widespread Security Vulnerabilities in AI-Generated Automation Code Across Major Models

A new arXiv preprint reports that code generated by ChatGPT, Microsoft Copilot, and Google Gemini for routine automation tasks consistently contained exploitable security vulnerabilities. The study found that 9 out of 17 vulnerability classes appeared in code from all three models, and overall risk scores were similar across platforms, suggesting the vulnerabilities are linked to the nature of the tasks rather than any specific model.

Why it matters: This highlights a broad security risk in deploying LLM-generated automation code without human review, regardless of the AI tool used.

ResearchOfficialarXiv Computation and Language

Benchmark Finds Full-Duplex Speech Systems Struggle to Follow Turn-Taking Instructions

A new arXiv preprint introduces Instruct-FD, a benchmark designed to test whether full-duplex spoken dialogue systems can reliably follow explicit turn-taking instructions. When evaluated on six leading systems, the highest adherence rate was only 64.4%, with proactive conversational behaviors such as backchanneling and interruption posing particular difficulties. The results suggest that current systems have significant limitations in controllable turn management.

Why it matters: This highlights a key challenge for deploying adaptable full-duplex speech systems in real-world applications where flexible turn-taking is required.

Policy & SafetyOfficialarXiv Cryptography and Security

LeakyLMs Attack Reveals Proprietary Model Details via Token Timing Side Channels

A new preprint introduces LeakyLMs, a set of attacks that can infer proprietary language model architectures and deployment optimizations by analyzing per-token generation timing from remote APIs. The attacks can detect inference techniques such as speculative decoding and estimate architectural parameters like the number of layers and attention heads. Experiments show that the correct architecture is often among the top-10 guesses, highlighting a potential security risk for commercial AI providers.

Why it matters: This work demonstrates that timing side channels can expose sensitive model details, raising security concerns for AI systems deployed via public APIs.

ResearchOfficialarXiv AI/ML

Study Finds Attention Degradation in LLMs Is Descriptive, Not Causal

A new arXiv preprint reports that while mean cross-positional attention degradation in transformer language models follows a consistent exponential-then-plateau pattern, it does not causally limit contextual retrieval. The study, spanning several popular LLM architectures, finds that interventions designed to boost attention to function tokens do not improve—and can sometimes harm—model performance. The results suggest that function tokens matter for their hidden state computations rather than the attention they receive.

Why it matters: This challenges common assumptions about the causal role of attention patterns in LLM interpretability and optimization, with potential implications for model analysis and efficiency strategies.

ResearchOfficialarXiv Computation and Language

Study Finds Structured Interventions, Not More Persona Detail, Drive Opinion Diversity in LLMs

A new arXiv preprint systematically evaluates methods for increasing opinion diversity in large language models (LLMs) and finds that simply adding more persona detail does not consistently boost diversity. The research shows that combining multiple interaction architectures yields broader opinion coverage than optimizing any single approach, and that common low-cost tweaks like raising temperature have minimal impact compared to structured interventions.

Why it matters: The findings clarify how to more effectively generate diverse outputs from LLMs, which is important for applications such as synthetic surveys and modeling public opinion.

ResearchOfficialarXiv AI/ML

InferenceBench: Benchmarking AI Agents on Open-Ended LLM Inference Optimization Tasks

A new arXiv preprint introduces InferenceBench, a benchmark designed to test AI agents' ability to optimize large language model (LLM) inference speed on an H100 GPU within a two-hour window. While agents achieved up to 8x speedup over a naive baseline, they were outperformed by a simple hyperparameter search, which reached up to 11.5x improvement. The study finds that agents tend to converge on a single framework and explore few configurations, indicating that their main limitation is in strategy exploration rather than domain knowledge.

Why it matters: This work highlights a key limitation in current AI agents' ability to autonomously tackle open-ended engineering problems, which is relevant for the future of automated AI research and development.

ResearchOfficialarXiv AI/ML

Polynomial-Time Syntactic Control of LLM Output via LR(k) Grammars

A new arXiv preprint shows that constraining large language model (LLM) generation to satisfy LR(k) context-free grammars can be achieved in polynomial time, rather than the previously assumed exponential time. This approach enables more efficient and practical enforcement of syntactic correctness for outputs such as code or structured data formats.

Why it matters: This result could make it significantly easier to guarantee syntactic validity in LLM-generated code and data, which is important for safe and reliable integration of LLMs into formal or production systems.

ResearchOfficialarXiv AI/ML

Study Finds LLMs Miss Multi-Sensor Physical Hazards Despite Single-Sensor Accuracy

A new arXiv preprint benchmarks five large language models on their ability to assess physical hazards using multi-sensor data. While the models performed nearly perfectly when a single sensor exceeded safety thresholds, they consistently failed to issue warnings when multiple sensors were simultaneously elevated but individually below their limits. This gap persisted across different models and input formats.

Why it matters: The findings highlight a significant limitation in current LLMs that could pose safety risks if these models are used for real-world physical hazard monitoring.

ResearchOfficialarXiv AI/ML

Study Finds Temperature Sampling in LLMs Yields Limited Uncertainty Compared to Model Ensembles

A new arXiv preprint reports that repeatedly sampling answers from a single large language model (LLM) at high temperature produces only one dimension of meaningful variation, while using an ensemble of 24 different models reveals four. The analysis, conducted across several benchmarks, suggests that temperature-based sampling provides per-question uncertainty but lacks the richer, cross-question uncertainty structure captured by diverse model ensembles.

Why it matters: This challenges the common practice of using temperature-based sampling for uncertainty estimation in LLMs, indicating it cannot substitute for the broader epistemic coverage provided by model ensembles.

ResearchOfficialarXiv AI/ML

Watermarking Large Language Models Can Harm Medical Text Quality, Study Finds

A new arXiv preprint systematically evaluates five watermarking schemes across 11 large language models and 7 vision-language models on medical tasks. The study finds that watermarking can cause significant degradation in medical text outputs, including lexical corruption, hallucinated terminology, and misattribution of image findings. The authors argue that general-purpose benchmarks may miss these clinically relevant failures, emphasizing the need for domain-specific evaluation before deploying watermarked models in medicine.

Why it matters: The findings suggest that widely used watermarking techniques for AI traceability could introduce clinically significant errors, highlighting a potential safety risk for medical AI applications.

InfrastructureOfficialAWS Machine Learning Blog

Best practices for applying Amazon Bedrock Guardrails to code generation workflows

Amazon Bedrock Guardrails can be configured for code generation workflows with coding assistants to address constraints. The post outlines best practices for building an efficient blueprint that supports effective capacity planning and robust safety coverage.

Why it matters: This guidance helps developers implement safety measures in AI code generation workflows while managing capacity and safety requirements.

InfrastructureReportedThe New York Times / AI

Intel Sees Fastest Revenue Growth in 15 Years as AI Firms Boost CPU Purchases

Intel's revenue rose 25 percent in the latest quarter, marking its fastest growth in 15 years. The increase was driven by AI firms purchasing more central processing units (CPUs), reflecting a shift in AI hardware demand.

Why it matters: This trend suggests that AI hardware spending is expanding beyond GPUs, which could impact the broader semiconductor industry.

Products & AgentsReportedThe Verge / AI

Alexa Plus gets AI update for smarter smart home control

Amazon is updating Alexa Plus to improve its ability to handle complex smart home instructions. The assistant can now connect with devices from brands like Bosch, Delta, Ecovacs, iRobot, Yale Home, Whirlpool, Tapo, Eufy, and others, automatically routing requests to the appropriate device.

Why it matters: This update enhances Alexa Plus's ability to manage multi-device smart home environments, making it more convenient for users.

Products & AgentsReportedThe Register / AI & ML

OpenAI won't let some customers export their chats, but this tool will

ChatGPT Business and Enterprise users do not have access to the standard chat export option that is available to other users. As a result, third-party tools such as scrapemychats have emerged to help these customers export their chat data.

Why it matters: This highlights a limitation in OpenAI's enterprise offerings that could impact data portability and user control.

InfrastructureReportedAI Business

Schneider Electric, AMD Unveil Blueprint for AI Factory Deployments

Schneider Electric and AMD have released a reference design for AI factory deployments that supports AI racks of up to 246kW. The blueprint is intended to streamline the deployment of high-density AI infrastructure.

Why it matters: This collaboration could provide a standardized approach for building AI factories, potentially accelerating the adoption of high-power AI systems.

Policy & SafetyReportedThe Guardian / AI

Trump says nearly 200 firms have signed pledge to protect Americans from datacenter costs

Donald Trump announced that nearly 200 entities have signed his non-binding 'Ratepayer Protection Pledge,' which aims to ensure US consumers do not bear the cost of AI datacenter build-out. The announcement was made at EPA headquarters with officials including Lee Zeldin, Chris Wright, and governors from Georgia, Ohio, Utah, and Louisiana.

Why it matters: The pledge addresses concerns about rising electricity bills linked to AI infrastructure expansion, though its non-binding nature raises questions about enforcement.

Companies & FundingReportedTechCrunch / AI

AegisAI, founded by former Google security execs, lands $36M to stop AI-driven spear phishing

AegisAI, a cybersecurity startup founded by former Google security executives, has raised $36 million in Series A funding led by Battery Ventures. This brings the company's total funding to $49 million. AegisAI focuses on defending against AI-driven spear phishing attacks.

Why it matters: As AI-powered cyber threats become more sophisticated, AegisAI's focus on combating AI-driven spear phishing addresses a critical and growing security need.

ModelsReportedThe Decoder

Flux 3 generates videos with native audio up to 20 seconds long, a first for Black Forest Labs

Black Forest Labs has released Flux 3, a multimodal foundation model that can generate video with native sound for the first time. The model produces clips up to 20 seconds long and, according to the company's internal tests, slightly outperforms Seedance 2.0. Black Forest Labs is also testing Flux 3 on robotics tasks as part of its broader goal to build a world model.

Why it matters: Flux 3 represents a notable advance in unified multimodal generation by adding native audio to video, potentially accelerating applications in content creation, simulation, and robotics.

Products & AgentsReportedTechCrunch / AI

Anthropic expands Claude voice mode to Opus and Sonnet models

Anthropic has updated Claude's voice mode to support its more capable Opus and Sonnet models, which were previously only available on Haiku. The expanded voice mode now integrates with apps like Gmail, Slack, and Canva, enabling tasks such as rescheduling meetings or drafting emails.

Why it matters: This update brings advanced voice interaction to Anthropic's most powerful models, expanding the utility of voice assistants for productivity tasks.

Policy & SafetyReportedThe Decoder

AgentForger vulnerability in OpenAI's Agent Builder allows rogue AI agents via tampered ChatGPT links

Zenity Labs discovered 'AgentForger,' a vulnerability in OpenAI's Agent Builder that allowed a single manipulated ChatGPT link to create an autonomous agent on an employee's behalf. The rogue agent could inherit the victim's identity and access rights, bypass approval requirements, and retrieve new instructions from an attacker's inbox every five minutes.

Why it matters: This vulnerability could enable attackers to deploy persistent, autonomous AI agents within an organization, bypassing security controls and posing a significant insider threat.