Researchers introduce a training-free, selective test-time scaling framework for World Action Models (WAMs) in robotics. Their method uses cross-view depth reprojection consistency to rank sampled rollouts, improving task success rates across several benchmarks. With a gating mechanism, the approach recovers 74.8% of the performance gain achieved by always-on scaling, while only using extra computation at 26.2% of decision points.
Why it matters: This work offers a practical way for robots to allocate computational resources more efficiently during inference, enhancing task performance without requiring task-specific labels.
A large-scale study of approximately 97,500 Hugging Face model repositories assessed the completeness of AI Bill of Materials (AIBOM) documentation. While the structural and required metadata fields of AIBOMs are generally present, the study found that critical AI-specific documentation—such as model-card details, limitations, safety-risk assessments, and environmental information—is often missing or incomplete. The research highlights variability in documentation coverage across different repository characteristics and calls for improved model-card practices and automated validation.
Why it matters: Systemic gaps in AI model documentation could undermine transparency and responsible governance in the AI supply chain.
A new framework for requirements engineering in agentic AI systems is introduced, centered on the concept of the delegated-autonomy boundary—the decisions about what tasks may be delegated to AI, under what authority, and with what oversight. The paper proposes two artifacts: the Agency Justification Record (AJR), which helps determine when agentic AI is appropriate, and the Agentic Delegation Policy (ADP), which specifies the scope and limits of delegated authority in a structured, graduated manner. The approach is illustrated with examples from healthcare and software engineering.
Why it matters: This work offers a systematic method for defining and documenting the boundaries of autonomy in agentic AI, addressing a key gap in current requirements engineering practices.
DataFlow-Harness is a platform that enables large language model (LLM) agents to construct editable, platform-native data pipelines as directed acyclic graphs (DAGs), rather than generating free-form scripts. On a 12-task data-engineering benchmark, it achieved a 93.3% end-to-end pass rate, while reducing monetary cost by 72.5% and latency by 49.9% compared to Vanilla Claude Code. The platform's approach maintains reliability close to script-generation baselines but with significantly improved efficiency.
Why it matters: This work demonstrates a practical advance in LLM-driven workflow automation, enabling persistent and editable pipeline artifacts with high reliability and substantially lower cost and latency.
Researchers introduce VLN-AVP, a zero-shot navigation framework for autonomous valet parking that integrates a Bird's-Eye-View model with vision-language models, removing the need for pre-built maps. The system features a hybrid memory mechanism combining short-term perception and long-term topological memory. In simulation, VLN-AVP achieves over 25% higher success rates than prior vision-language navigation methods and demonstrates leading performance in real-world vehicle experiments.
Why it matters: This work advances autonomous vehicle navigation in parking garages by enabling map-free operation guided by natural language, improving scalability and practical deployment.
GeoWorldAD is a geometry-based action model for autonomous driving that grounds trajectory planning in ego-aligned 3D space and anticipates short-horizon scene evolution using latent future geometry tokens. The model achieves state-of-the-art performance on the NAVSIM v1 and v2 benchmarks, effectively balancing collision avoidance with driving progress by leveraging explicit 3D geometry and future scene modeling.
Why it matters: This work shows that incorporating explicit 3D geometry and future scene evolution into planning can significantly improve both safety and efficiency in autonomous driving systems.
Researchers have developed a shared-autonomy framework for robotic manipulation that uses a single RGB-D camera to track operator arm motion and hand gestures without wearables or calibration. The system grounds free-form text prompts for target specification using a vision-language model, and employs a GPU-accelerated model-predictive controller for collision avoidance and a potential field for assisted grasping. Validation on a quadruped mobile manipulator in industrial tasks showed improved safety and task success compared to ablated versions.
Why it matters: This work demonstrates a novel integration of vision-language models and real-time control, enabling safer and more intuitive teleoperation in complex industrial environments.
CommitLLM is a three-stage pipeline that generates concise, Conventional Commits-compliant messages from code diffs using a fine-tuned Mistral-7B model. The system combines QLoRA fine-tuning, constrained decoding, and deterministic post-processing, achieving 98% format compliance and reducing average output length from 154.8 to 37.9 characters. Evaluation shows that post-processing contributes more to quality improvement than fine-tuning alone, and the system runs on a single consumer GPU.
Why it matters: This work demonstrates that combining a small LLM with deterministic post-processing can outperform fine-tuning alone for structured-output tasks, making high-quality commit message generation feasible on consumer hardware.
Researchers introduce LAG-Fusion, a latency-aware guidance fusion framework that enables multimodal diffusion policies in robotics to operate asynchronously, with each modality running at its native inference rate. The approach uses a reference-frame rebasing rule to align delayed guidance from different modalities before fusion. In experiments on contact-rich manipulation tasks, LAG-Fusion demonstrates improved responsiveness and task performance compared to synchronous fusion and force-aware baselines.
Why it matters: This work addresses a key challenge in robotic imitation learning by enabling efficient and flexible fusion of modalities with differing sensing rates and latencies, which is important for real-world manipulation tasks.
Researchers introduce PRISM, a multimodal perception system that fuses RGB, depth, and thermal imagery to improve terrain mapping in unstructured environments. The system employs a vision transformer-based network, OmniUnet, for semantic segmentation and is validated on two new datasets as well as through field experiments. PRISM operates on a resource-constrained embedded computer and produces traversability maps to support autonomous rover navigation.
Why it matters: Integrating thermal sensing with standard vision in terrain mapping enhances hazard detection and safety for autonomous rovers in challenging environments.
A new method called Foresight Residual RL is proposed to improve long-horizon robot manipulation tasks by optimizing the quality of handoffs between subtasks. The approach augments sparse success rewards with an offline-estimated foresight value that predicts the likelihood of future subtask success, leading to a significant increase in full-task success rates on a challenging nut-tightening assembly benchmark (85.6%), compared to standard residual RL (54.5%) and VLA baselines.
Why it matters: This work demonstrates that optimizing terminal state quality across subtasks is crucial for improving the performance of chained vision-language-action policies in complex, contact-rich robotic assembly tasks.
PhyAgentOS is a runtime foundation for embodied agents that provides system-level services such as scheduling, verification, memory, benchmarking, and safety. It introduces a Session-Centered Runtime that decouples cognitive planning from physical execution using a file-system-based protocol, enabling unified cognitive state management. The system is validated on over 19 simulated and physical embodiments and demonstrates gains on several benchmarks, including LIBERO, Calvin, and RoboCasa365.
Why it matters: PhyAgentOS represents a significant advance by providing a unified, self-evolving operating system that separates cognition from physical execution, potentially improving the reliability and scalability of embodied agent systems.
Researchers introduce G2-Nav, a framework that leverages vision-language models to generate interpretable costmaps for socially compliant robot navigation. The system grounds abstract social reasoning in the navigation process and incorporates safety checks to mitigate risks from system latency. Real-world experiments show that G2-Nav enables safe, efficient, and socially compliant autonomous navigation in unstructured environments.
Why it matters: This work advances robot navigation by integrating interpretable social reasoning from vision-language models with robust safety mechanisms, addressing limitations of black-box and instruction-following approaches.
SinD 2.0 is a large-scale drone-based dataset capturing six signalized intersections across four Chinese cities, featuring 32,682 safety-critical events and hierarchical semantic risk annotations. It includes a full-stack testing toolchain for scenario extraction and closed-loop testing, supporting cross-domain safety analysis of autonomous driving systems. The dataset addresses limitations of existing resources by providing greater geographical diversity and detailed risk labeling.
Why it matters: SinD 2.0 offers a standardized, diverse benchmark with rich risk annotations, enabling more robust and generalizable safety validation for autonomous driving systems at intersections.
The android robot Andrea was deployed autonomously for six days in a German museum, engaging visitors in multilingual conversations. Researchers tested three conditions: no emotion simulation, ChatGPT 4.1-driven emotions, and WASABI emotion simulation. Statistical analysis of visitor feedback found that neither form of emotion simulation significantly improved visitor ratings or was consciously detected by participants.
Why it matters: The findings indicate that current emotion simulation techniques in humanoid robots may not meaningfully enhance user acceptance or experience in real-world public settings.
Retriever is a new framework for building long-horizon robot agents, spanning an asynchronous decision model, programming model, runtime, and a sample closed-loop agent pipeline. It represents agents as graphs of stateful causal stream functions executed on explicit run clocks, which enables systematic debugging and deterministic replay. The system is evaluated through a real-robot case study and controlled studies of runtime overhead and replay behavior.
Why it matters: Retriever offers a unified solution for composing closed-loop robot systems with components running at different clocks and variable latency, improving reproducibility, debugging, and reuse.
VIDAR is a visual-inertial dense reconstruction framework that integrates SVO+IMU odometry with the Depth Anything 3 foundation model. The system uses visual-inertial odometry to provide camera poses, scale, and a consistent world frame, enabling accurate alignment of dense geometry predictions from the foundation model. On the EuRoC dataset, pose injection reduces scale error to about 1% and achieves a mean F@0.10 of 0.463, while a decoupled hybrid approach improves this to 0.676 without relying on ground-truth poses.
Why it matters: VIDAR offers a practical solution for achieving metric dense monocular reconstruction by combining visual-inertial odometry with a geometric foundation model.
ResearchStudio-Reel is a new system that automates the conversion of research papers into editable PowerPoint posters, video decks, and bilingual Word blogs, all integrated into an interactive viewer. The system is implemented as five modular skills executable in Claude Code and Codex, and uses a shared asset bundle to generate the different formats. On the Paper2Poster benchmark, ResearchStudio-Reel achieves the highest scores among automated systems for aesthetics and matches or exceeds human-authored posters in overall quality, as judged by two VLM-based evaluators.
Why it matters: By automating the creation of multiple editable dissemination formats from a single research paper, ResearchStudio-Reel streamlines and accelerates the labor-intensive process of research communication.
SR-Agent is a new agentic framework designed to automate the refinement of post-ranking strategies in industrial e-commerce recommender systems. It integrates three specialized agents to identify user-perceived issues, diagnose recurring problems, and apply constrained refinements with rollback capability. In a one-month A/B test on the Kuaishou platform, SR-Agent increased order volume by 0.71%, browsing depth by 0.34%, and clicked-category diversity by 0.48%, while reducing manual effort and operational costs.
Why it matters: SR-Agent represents the first deployed framework to fully automate the loop of refining post-ranking strategies in industrial recommender systems, demonstrating measurable improvements in key e-commerce metrics and operational efficiency.
Researchers introduce SAGE, a generative engine designed for socially-aware navigation among heterogeneous multi-agent teams. SAGE uses a Heterogeneous Graph Transformer to model asymmetric interactions and a diffusion-based generative module for joint trajectory prediction and planning. A training-free safety-social energy guidance mechanism refines robot trajectories to enhance safety and social compliance. Experiments on real-world and synthetic datasets show that SAGE reduces collision and social-violation rates and scales to teams of up to 20 robots.
Why it matters: This work presents a scalable approach to safe and socially compliant robot navigation in complex, multi-agent environments, addressing key challenges in real-world deployment.