Researchers have established the first non-vacuous generalization bounds for parameter-efficient reinforcement learning with verifiable rewards (RLVR) fine-tuning at the billion-parameter scale. They introduce Progressive RLVR, a framework that integrates RLVR with on-policy distillation, TinyLoRA, and model quantization, retaining 84-97% of standard LoRA performance while producing models that are 14,796x more compressible. The resulting generalization bounds are within 6-11% of fine-tuned model accuracy across four domains: mathematical problem-solving, programming, general-knowledge reasoning, and Text-to-SQL.
Why it matters: This work provides the first theoretical generalization guarantees for RLVR fine-tuning at scale, bridging a critical gap between practical application and theoretical understanding in large language model training.
Researchers have introduced an active evaluation framework that models policy testing as a sequential experimental design problem. By using a probabilistic surrogate model to adaptively select test configurations, the method enables more efficient and systematic evaluation of generalist robot policies. In experiments involving 2331 real-world trials across three tasks and three factor variations, the framework reduced the number of required trials by 20-40% compared to random testing.
Why it matters: This approach addresses a major bottleneck in robotics by making real-world evaluation of generalist robot policies more sample-efficient and comprehensive.
Researchers introduce xHC (Expanded Hyper-Connections), a method that enables Transformer models to expand their residual streams beyond the previous limit of N=4. By combining temporal feature augmentation and a sparse residual-stream architecture, xHC achieves strong and consistent downstream improvements in 18B and 28B MoE models. The xHC-Flash variant further reduces memory traffic, making large-N residual-stream expansion practical for large language model pre-training.
Why it matters: This work establishes a new, practical scaling axis for large language models, enabling more efficient pre-training and consistent performance gains beyond traditional width and depth scaling.
Researchers have developed COAT (Counterfactual Optimal Action Tree), a framework that learns interpretable prescriptive policies from observational data by combining counterfactual outcome estimation with mixed-integer optimization. In a 17-week field pilot with a major global airline, COAT increased upsell revenue per booking by 6.9%, with the airline projecting $50–$150 million in incremental annual premium seat revenue. The pilot's success led to scaled adoption and influenced broader AI-driven decision initiatives within the organization.
Why it matters: COAT provides a practical, transparent approach for deriving actionable business policies from observational data, demonstrating significant real-world financial impact in the airline industry.
Researchers introduce TEDDY, a 1.84-million-parameter decoder transformer trained on 73 million ICD-10 diagnoses from 1.6 million children at a single pediatric institution. TEDDY achieved a median AUC of 72.0% across 797 disease-onset prediction tasks, outperforming several larger and commonly used baseline models. The model demonstrated strong performance even for rare diseases and could detect predictive signals more than two years before diagnosis.
Why it matters: This work shows that compact generative models can provide accurate, early risk predictions for a wide range of pediatric diseases, including rare conditions, without requiring massive datasets or very large models.
Researchers introduce RENEW, a method that leverages human preferences over imagined rollouts to address model exploitation in offline reinforcement learning. The approach, formalized as Dynamics Learning from Human Feedback (DLHF), uses human intuition to identify unrealistic model dynamics. RENEW improves the practicality of preference-based supervision by focusing finetuning on regions where the model is most uncertain, thereby enhancing sample efficiency and reducing exploitation.
Why it matters: This work presents a novel approach to mitigating model exploitation in offline RL by directly incorporating human feedback, potentially reducing reliance on costly expert demonstrations.
A new preprint introduces Nous, a belief-based memory architecture for LLM agents that uses Bayesian inference and surprise-driven updates. The study finds that belief updating offers little benefit over simpler methods on standard benchmarks, but significantly outperforms them when observations vary in trustworthiness. The authors also propose provenance-capped updating to defend against memory poisoning attacks, and highlight evaluation discrepancies in long-term memory benchmarks.
Why it matters: This work clarifies when probabilistic memory is genuinely useful for LLM agents and introduces practical defenses against memory poisoning, which is important for deploying reliable agents in adversarial settings.
Researchers have introduced LIGO-PINN, a new framework that addresses convergence failures in physics-informed neural networks (PINNs) by optimizing network weight initialization. In evaluations on challenging partial differential equation (PDE) domains—including 2D fluid dynamics and 3D unstructured domains—LIGO-PINN achieved an average performance improvement of 91.5% over six baselines and 81% over the strongest baseline. The method demonstrates improved reliability and generalization in PINN training without relying on expensive hyperparameter tuning or complex training strategies.
Why it matters: This work offers a significant advance in the robustness and effectiveness of PINNs by targeting weight initialization, a previously underexplored factor, potentially broadening the practical applicability of PINNs in scientific and engineering domains.
A new method called Dysco is proposed to address instability in federated fine-tuning of large models using Low-Rank Adaptation (LoRA). Dysco dynamically allocates client-specific subspaces to reduce data-parameter interference, leading to more stable aggregation of updates. Experiments demonstrate up to a 9x reduction in synthetic training loss and up to a 4.3% improvement in clinical-note classification tasks compared to baselines. The approach also maintains low computational overhead and outperforms recent federated LoRA methods.
Why it matters: This work offers a significant advance in federated learning by mitigating a key source of instability in LoRA-based fine-tuning, enabling more reliable and accurate collaborative model training across diverse clients.
SearchOS-V1 is a multi-agent framework designed to improve open-domain information-seeking by externalizing search progress into an explicit, persistent, and shared state. The system introduces Search-Oriented Context Management (SOCM) and a pipeline-parallel scheduling mechanism, which together help prevent repetitive search loops and improve agent collaboration. SearchOS-V1 achieves state-of-the-art results on the WideSearch and GISA benchmarks, outperforming existing single- and multi-agent baselines.
Why it matters: By making search progress explicit and shared, SearchOS-V1 addresses a key limitation of current AI agents—repetitive search loops—potentially increasing the reliability and efficiency of autonomous information-seeking systems.
A new preprint presents the first comprehensive study of how fairness-enhancing algorithms impact membership inference privacy risks at the subpopulation level. The authors adapt the Likelihood Ratio Attack for subgroup auditing, revealing privacy disparities that aggregate evaluations can obscure. They also show that the benefits and costs of differential privacy are unevenly distributed across subpopulations, and introduce a unified empirical framework for jointly evaluating fairness, privacy, and utility.
Why it matters: This work demonstrates that fairness interventions can have uneven privacy impacts across subpopulations, underscoring the need for subgroup-level auditing to avoid hidden disparities.
Researchers introduce spatiotemporal augmentations that simulate common streaming artifacts—such as pixelation, blur, and ghosting—to train imitation learning agents for 3D video games. Agents trained with these augmentations achieve up to 41% higher performance under stable streaming conditions and show much less performance degradation (7.45% vs 49.82%) under network lag compared to agents trained without such augmentations.
Why it matters: This approach provides a practical and data-efficient way to make game-playing AI more robust to real-world streaming conditions, narrowing the gap between training and deployment.
Researchers have introduced a framework that uses a lightweight router to dynamically decide, for each instance, whether a large language model should apply reasoning or direct inference in ranking tasks. This approach improves ranking accuracy while significantly reducing token usage. Experiments on the MovieLens dataset with Qwen3-4B show up to a 6.3% gain in NDCG@10 and a 49.5% reduction in token consumption.
Why it matters: This work offers a practical method to balance accuracy and computational efficiency in LLM-based retrieval and recommendation systems by allocating reasoning only when it is likely to be beneficial.
ShopX is a foundation model designed to unify intent understanding, execution planning, and item-space operations for agentic shopping. By leveraging semantic IDs, ShopX enables direct translation of complex user intents into item-space outcomes, reducing information loss between agent orchestration and item fulfillment. Evaluation on tasks derived from Taobao production logs indicates that ShopX improves fulfillment behavior, particularly for complex or ambiguous requests.
Why it matters: ShopX represents a notable advance in bridging language-based shopping agents and item fulfillment, enabling more effective and natural intent-driven shopping experiences.
A new preprint introduces the sublinear-growth principle for deep residual architectures, identifying a sharp stability threshold based on the input-magnitude exponent (q ≤ 1) of residual blocks. The authors provide theoretical arguments showing this criterion is both necessary and sufficient for stable training, clarifying the stabilizing role of layer normalization and enabling efficient certification of architectural stability. The work also demonstrates a parameter-free modification that stabilizes the Mamba block without normalization, with experiments confirming stable training for q ≤ 1 variants.
Why it matters: This research offers a principled, theoretically grounded method for designing and certifying stable deep residual networks, potentially reducing reliance on empirical trial-and-error in architecture development.
Researchers propose gate-zero growth, a function-preserving operator for continual learning that adds new residual blocks to neural networks via zero-initialized gates. In experiments with a 300M-to-857M parameter Transformer transitioning from WikiText-103 to BookCorpus, gate-zero growth achieves near-zero forgetting, while a non-function-preserving control exhibits significantly higher forgetting. The framework also unifies the geometric analysis of related methods such as LoRA, ReZero, and zero-init adapters.
Why it matters: This work offers a principled approach to expanding neural network capacity without catastrophic forgetting, which is important for scalable continual learning in large models.
Researchers introduce Branching Policy Optimization (BPO), a reinforcement learning algorithm designed for large language model (LLM) agents operating in deterministic, snapshottable sandboxes. BPO leverages the ability to fork alternative actions at high-entropy decision points, sharing rollout prefixes to reduce variance and compute unbiased, lower-variance advantage estimates from sibling returns. Experiments on WebShop, ALFWorld, and SWE-bench Verified show that BPO improves success rates by 3.6–6.1 absolute points over GRPO and RLOO at matched compute, and achieves similar performance to the best baseline with 38% fewer policy updates.
Why it matters: BPO demonstrates a novel and more efficient approach to RL for LLM agents by exploiting sandbox determinism, leading to improved sample efficiency and training stability.
Researchers introduce CoSimRec, an offline agent-based evaluation framework designed to model the interplay between coordinated accounts, dynamic ranking, and user responses in recommender system feedback loops. The framework proposes the Algorithmic Penetration Rate (APR) metric family to quantify the extent to which target content reaches and engages non-bot users. Experiments across multiple datasets show that popularity-based and feedback-sensitive ranking algorithms can significantly increase coordinated content penetration, while synchronization-aware ranking strategies can mitigate this effect.
Why it matters: This work offers a systematic approach to evaluating how recommender systems may amplify coordinated content, addressing a critical gap in robustness assessments.
A new preprint proposes that memory for coding agents should be integrated directly into Git version control, rather than relying on separate retrieval systems. The authors demonstrate that this approach achieves a pooled MRR of ~0.31 for seed supply and 0.83 answer sufficiency on a production system, with results that are replicable at zero labeling cost. The system leverages commit-session links to provide ground truth, enabling practical and scalable memory for agentic development workflows.
Why it matters: This work introduces a novel, practical method for agentic development memory by leveraging existing Git infrastructure, potentially improving reproducibility and efficiency in coding agent workflows.
Researchers have introduced AE-UAV, the first airborne-captured event camera dataset specifically for air-to-air UAV tracking, featuring 178 flight sequences with detailed annotations. They also present FSFT, a lightweight, training-free tracker that achieves 420 FPS on CPU-only hardware and retains 93.97% of the accuracy of state-of-the-art GPU-based methods. This approach offers a 5.32-fold speedup and demonstrates strong generalization in temporal resolution.
Why it matters: This work enables efficient, real-time UAV tracking on resource-constrained platforms, addressing a major challenge in airborne remote sensing.