Enterprise AI news — Page 2

How businesses are adopting AI through enterprise software, workplace tools, automation, and production deployments.

ResearchOfficialarXiv Software Engineering

Where Does Agent Reliability Come From? A Cross-Benchmark Decomposition of Verification Loops, Specialist Models, and Scaffolding in a Production Enterprise Agent

A preprint analyzes the Leni production enterprise agent to determine the sources of reliability improvements across three public benchmarks. The study finds that while the full system outperforms its base model by up to 15 percentage points, most of the reliability gains come from scaffolding and specialist models, with the verification loop itself contributing a smaller but targeted improvement (+1.5 points), especially for the most challenging tasks. The work provides a detailed decomposition of these effects and presents empirical data on the verification loop's performance.

Why it matters: This decomposition clarifies which architectural components most effectively boost reliability in enterprise agent systems, informing future design choices for robust multi-step agents.

ResearchOfficialarXiv Multiagent Systems

Agentic ERP: Multi-Agent LLM Architecture for Autonomous Enterprise Resource Planning

A new preprint introduces Agentic ERP, a multi-agent architecture that leverages role-aligned large language model (LLM) agents and human oversight to automate operational decision-making in enterprise resource planning (ERP) systems. The system employs a graph-based orchestrator and was evaluated through a 365-day simulation, where it achieved zero stockouts, outperforming rule-based baselines that experienced hundreds of stockouts. The evaluation also included scenario-based tasks and crisis management comparisons, demonstrating significant improvements over traditional approaches.

Why it matters: This work provides evidence that LLM-based multi-agent systems, when combined with human oversight, can enable ERP systems to autonomously execute operational decisions, potentially advancing enterprise automation beyond traditional rule-based methods.

ResearchOfficialarXiv Information Retrieval

AnnoRetrieve: Structured Retrieval for Unstructured Documents

AnnoRetrieve introduces a new retrieval paradigm that replaces traditional vector embeddings with lightweight structured queries over automatically generated annotation schemas. Using SchemaBoot for schema induction and Structured Semantic Retrieval (SSR) for precise matching, AnnoRetrieve enables annotation-driven semantic retrieval without relying on LLM calls. Experiments on real-world datasets demonstrate that this approach significantly reduces LLM usage and retrieval costs while maintaining high accuracy.

Why it matters: AnnoRetrieve's annotation-driven approach could substantially reduce the computational and financial costs of large-scale document analysis, making precise retrieval more accessible and scalable.

ModelsOfficialarXiv Information Retrieval

WHALE: A Scalable Unified Model for Recommendation with Wukong-HSTU Architecture

The WHALE model unifies non-sequence and sequence feature modeling for recommendation systems by integrating Wukong and HSTU modules with an attention-based fusion mechanism. The architecture maintains both modules throughout the network, enabling high-order feature interactions to leverage detailed user behavior histories. WHALE demonstrates consistent improvements in offline experiments and delivers positive online gains in industrial settings, with deployment in production systems.

Why it matters: WHALE provides a practical and scalable approach to combining complementary recommendation architectures, showing real-world deployment and measurable improvements.

Policy & SafetyOfficialarXiv Cryptography and Security

ShadowPickle: Stealthy Pickle Deserialization Attacks Evade ML Model Scanners and Hubs

Researchers introduce ShadowPickle, a set of three novel pickle deserialization attacks that evade detection by ten state-of-the-art model scanners and four major model hubs, including Hugging Face. These attacks exploit the Pickle VM's external module import mechanism to execute malicious payloads during model deserialization. ShadowPickle achieves up to 63% evasion rates and up to 50% higher evasion rates than previous attacks, exposing significant weaknesses in current model scanning defenses.

Why it matters: This work reveals critical security vulnerabilities in widely used machine learning model hubs, raising concerns about supply chain attacks and the safety of shared pre-trained models.

ResearchOfficialarXiv Computation and Language

HALO: Hallucination-Aware Layered Oversight for Trustworthy Enterprise AI

A new preprint argues that achieving zero hallucination in large language models (LLMs) cannot be accomplished by the model alone, but must be enforced at the system level. The authors introduce HALO, a six-layer assurance architecture that includes grounded generation, constrained execution, multi-signal verification, calibrated abstention, total traceability, and continuous oversight. The framework is demonstrated on a regulated claims-extraction task.

Why it matters: This work reframes hallucination as a system-level failure mode and proposes a practical, multi-layered architecture for deploying trustworthy AI in regulated enterprise environments.

Products & AgentsOfficialAWS Machine Learning Blog

Build specialized agent workflows with Amazon Quick and NVIDIA NeMo Agent Toolkit

Amazon Quick and the NVIDIA NeMo Agent Toolkit can be combined to create specialized agent workflows for business users. An example in supply-chain risk demonstrates how users can move from an Amazon Quick dashboard to receiving guided mitigation recommendations.

Why it matters: This integration allows business users to access AI-driven agent workflows directly from dashboards, potentially streamlining decision-making processes.

ModelsOfficialarXiv Information Retrieval

RecGPT-V3: Stateful, Hybrid-Modal Recommender System Deployed on Taobao

RecGPT-V3 is a stateful, hybrid-modal recommender system deployed in Taobao's 'Guess What You Like' feed. It introduces a Memory Hub to reduce user-modeling computation by 55.8%, a Hybrid-modal Foundation Model for joint reasoning over text tags and Semantic IDs, and Latent Intent Reasoning to lower output token cost by 200x. Large-scale online A/B tests report improvements in IPV (+1.28%), CTR (+1.00%), TC (+1.97%), GMV (+3.97%), and a 52.4% reduction in serving resource consumption.

Why it matters: RecGPT-V3 demonstrates a significant advance in scaling LLM-based recommender systems, achieving notable gains in both user experience and resource efficiency in a real-world, high-traffic deployment.

ResearchOfficialarXiv Computation and Language

Robust Explanations for User Trust in Enterprise NLP Systems

A new preprint introduces a unified black-box robustness evaluation framework for token-level explanations in enterprise NLP systems. The study systematically compares encoder and decoder models, finding that decoder LLMs yield significantly more stable explanations (73% lower flip rates on average) and that explanation stability increases with model scale (44% improvement from 7B to 70B parameters). The work also presents a practical cost-robustness tradeoff curve to inform model selection before deployment in compliance-sensitive environments.

Why it matters: This framework addresses a key challenge in validating explanation robustness for black-box NLP systems, supporting user trust and regulatory compliance in enterprise applications.

ModelsOfficialarXiv AI/ML

Multi-Expert Consensus Framework Enhances Serverless Autoscaling with Cost and Dependency Awareness

A new autoscaling framework for serverless environments combines graph-based dependency analysis, short-term workload forecasting using multiple neural models (MLP, LSTM, CNN), and cost-aware scaling control. The approach uses a probabilistic ensemble of predictors, achieving 99.88% prediction accuracy and reducing infrastructure costs while maintaining performance targets in experiments with real workload traces. The framework also incorporates cold-start awareness and evaluates performance across multiple cloud pricing models.

Why it matters: This research demonstrates a robust and practical advance in serverless autoscaling, addressing key challenges of workload prediction, cost efficiency, and dependency management in cloud applications.

Companies & FundingOfficialOpenAI News

OpenAI CFO Sarah Friar Proposes Practical AI Scorecard

OpenAI CFO Sarah Friar has introduced a practical AI scorecard designed to measure return on investment (ROI) using metrics such as useful work, cost per successful task, dependability, and return on compute. The framework is intended to help organizations evaluate the effectiveness of their AI investments.

Why it matters: A standardized scorecard could help organizations make more informed decisions about adopting and scaling AI technologies.

ResearchOfficialarXiv Machine Learning

COAT: Interpretable Prescriptive Policies from Observational Data Boost Airline Revenue

Researchers have developed COAT (Counterfactual Optimal Action Tree), a framework that learns interpretable prescriptive policies from observational data by combining counterfactual outcome estimation with mixed-integer optimization. In a 17-week field pilot with a major global airline, COAT increased upsell revenue per booking by 6.9%, with the airline projecting $50–$150 million in incremental annual premium seat revenue. The pilot's success led to scaled adoption and influenced broader AI-driven decision initiatives within the organization.

Why it matters: COAT provides a practical, transparent approach for deriving actionable business policies from observational data, demonstrating significant real-world financial impact in the airline industry.

ResearchOfficialarXiv AI/ML

Intention Abstraction Layer Enables Pre-Execution Conflict Detection in Autonomous Industrial Systems

A new preprint introduces the Intention Abstraction Layer (IAL), a middleware that leverages large language models and OWL ontologies to represent human intentions as persistent runtime objects in autonomous industrial systems. The IAL enables pre-execution detection and explanation of goal conflicts between autonomous agents, as demonstrated in a proof-of-concept scenario where conflicting intentions were flagged before execution. This approach shifts behavioral assurance from post-hoc analysis to intention-level checking.

Why it matters: The IAL offers a novel method for improving safety and reliability in multi-agent industrial systems by identifying and explaining goal conflicts before they lead to operational failures.

ResearchReportedVentureBeat / AI

Enterprise AI faces a trust gap as agents produce confident wrong answers from unreliable context

A VentureBeat Pulse Research survey of 101 enterprises found that 57% have observed AI agents producing confident but incorrect answers due to missing or inconsistent business context in the past six months. Provider-native retrieval methods, such as OpenAI's file search, have surpassed dedicated vector databases in usage, while 58% of enterprises are building or running a governed semantic layer to address the trust gap.

Why it matters: This highlights that the main challenge for enterprise AI is ensuring trust in the context provided to agents, with most organizations still developing the necessary infrastructure for reliability.

InfrastructureReportedVentureBeat / AI

The AI compute gap: Enterprises are buying infrastructure faster than they can measure what it costs

A VentureBeat Pulse Research survey of 107 enterprises highlights a growing gap between AI infrastructure investment and the ability to track its economics. Most organizations currently run AI on hyperscalers and model APIs, but 45% plan to evaluate specialized AI clouds within the year, despite almost none using them today. GPU utilization is at 50% or less for 83% of respondents, and fewer than half (44%) rigorously track compute costs.

Why it matters: This compute gap means enterprises are spending heavily on AI infrastructure without the visibility to control costs or optimize utilization, risking inefficiency and budget overruns.

Companies & FundingReportedTechCrunch / AI

Applied Computing raises $20M to build AI foundation model for oil and gas plants

Applied Computing has raised a $20M Series A to develop a foundation AI model for the oil, gas, and petrochemical industry. The company aims to provide operators with an AI system designed for entire plant operations.

Why it matters: This funding highlights increasing investment in specialized AI models for heavy industry, which could improve operational efficiency in oil and gas.

ResearchOfficialarXiv Computers and Society

Persona Migration and Expectation Recalibration in Generative AI Adoption: A Longitudinal Study at a State Department of Transportation

A longitudinal study of Microsoft 365 Copilot adoption at a state Department of Transportation found that perceived usefulness of the tool declined significantly after an eight-week pilot, while perceived ease of use, behavioral intention, and trust showed only minor, nonsignificant changes. The research identified three baseline user personas—Skeptics, Cautiously Positive users, and Champions—with substantial individual movement between groups, including 68% of Champions shifting to less enthusiastic personas. The study also observed that while some concerns (accuracy, privacy) decreased, worries about job impact and required skills increased.

Why it matters: This study provides empirical evidence that employee enthusiasm for generative AI tools can wane after real-world use, highlighting the need for ongoing expectation management and tailored support in enterprise AI deployments.

ModelsOfficialarXiv Computer Vision

LPM: Industrial-Scale Generative Video Restoration

Kuaishou's Large Processing Model (LPM) is a diffusion-based generative framework for photorealistic video restoration, designed to handle diverse, real-world degradations in user-generated content. LPM is reportedly the first generative video restoration model deployed at industrial scale, processing videos that account for about 45% of total viewing time on Kuaishou. The system achieves a 20% bitrate reduction compared to Kuaishou's in-house codec, resulting in substantial annual bandwidth cost savings.

Why it matters: This work demonstrates the first industrial-scale deployment of generative video restoration, showing that such models can deliver practical, scalable, and cost-effective improvements in large-scale video processing.

Policy & SafetyOfficialarXiv Cryptography and Security

Adversarial Prompting Framework Systematically Evaluates AI Model Safety

A new preprint introduces an Adversarial Prompting Framework (APF) designed to systematically assess the safety of AI models against adversarial prompt attacks. The framework generates structured prompts at varying levels of sophistication, from straightforward harmful requests to advanced encoding-based attacks, and enables automated testing with quantitative security metrics. The study finds notable differences in model vulnerabilities, with encoded prompts most frequently bypassing safety mechanisms.

Why it matters: This framework provides a practical, automated approach for identifying and quantifying critical vulnerabilities in AI models, which is essential for improving AI safety.

ResearchReportedVentureBeat / AI

Enterprise AI agents are mostly chatbots, reveals VentureBeat Pulse Research

A VentureBeat survey of 101 enterprises found that 71% report a quarter or fewer of their deployed 'agents' are true multi-step orchestrated workflows, with most being single-prompt chatbot wrappers. Anthropic's Claude is the primary platform for 40% of enterprises, chosen for its model alignment and reliable multi-step execution.

Why it matters: The gap between enterprise ambitions for agentic AI and the current reality highlights a risk of investing in orchestration infrastructure before deploying genuine multi-step agents.