A new preprint audits Portugal's publicly funded 9B language model AMALIA, evaluating its ability to code moral foundations in European Portuguese. The study finds that while AMALIA matches much larger open models in agreement with human coders, only about half of its coding performance can be attributed to the explicit theory underlying the coding scheme. The authors introduce a 'recovery gap' method to assess whether LLMs genuinely measure theoretical constructs or rely on surface correlations, and show that a larger multilingual model closes this gap, implicating limitations in AMALIA itself.
Why it matters: This work questions the epistemic trustworthiness of sovereign language models and introduces a portable audit method for evaluating their validity as scientific instruments.
AnnoRetrieve introduces a new retrieval paradigm that replaces traditional vector embeddings with lightweight structured queries over automatically generated annotation schemas. Using SchemaBoot for schema induction and Structured Semantic Retrieval (SSR) for precise matching, AnnoRetrieve enables annotation-driven semantic retrieval without relying on LLM calls. Experiments on real-world datasets demonstrate that this approach significantly reduces LLM usage and retrieval costs while maintaining high accuracy.
Why it matters: AnnoRetrieve's annotation-driven approach could substantially reduce the computational and financial costs of large-scale document analysis, making precise retrieval more accessible and scalable.
Research→Official→arXiv Audio and Speech Processing
ChipChat introduces a novel low-latency cascaded system for real-time on-device voice agents, integrating streaming speech recognition, large language models, text-to-speech, vocoder, and speaker modeling. Implemented in MLX, the system achieves sub-second response latency on a Mac Studio without dedicated GPUs, enabling privacy-preserving, fully on-device processing. The work demonstrates that architectural innovations and streaming optimizations can overcome traditional latency bottlenecks in cascaded systems.
Why it matters: This research provides a practical solution for real-time, privacy-preserving voice-based AI agents on consumer hardware, addressing a key challenge in deploying conversational AI locally.
Research→Official→arXiv Audio and Speech Processing
A preprint introduces Precision-Varying Prediction (PVP), a method that improves the adversarial robustness of automatic speech recognition (ASR) models by randomly varying the inference precision. The approach also enables adversarial example detection using a Gaussian classifier on outputs from different precision levels. When combined with uncertainty-based defenses, PVP increases the difficulty for adaptive attackers, requiring them to introduce more perceptible noise to evade detection. Experimental results show significant improvements in robustness and detection across multiple ASR models, languages, and attack types.
Why it matters: This work proposes a practical and effective defense against adversarial attacks on ASR systems, which are increasingly used in real-world automated applications.
Research→Official→arXiv Audio and Speech Processing
Researchers present X-Translator, a modular, low-cost, open-source speech-to-speech translation system that integrates streaming automatic speech recognition (ASR), machine translation, and prompt-conditioned text-to-speech (TTS). The system is designed for real-time translation in long-form and multi-speaker conversations, addressing challenges such as unstable ASR hypotheses, ambiguous turn boundaries, and speaker consistency. X-Translator is evaluated on translation quality, speech quality, latency, and speaker preservation, with code and a demo publicly available.
Why it matters: X-Translator offers an open, reproducible platform for real-time multilingual speech translation with speaker awareness, helping advance practical deployment in complex conversational scenarios.
Research→Official→arXiv Audio and Speech Processing
Researchers introduce WildElder, a Mandarin elderly speech corpus collected from online videos and annotated with transcription, speaker age, gender, and accent strength. The dataset addresses the scarcity of diverse, real-world elderly speech data for automatic speech recognition and speaker profiling. Experimental results demonstrate the challenges of elderly speech recognition and establish WildElder as a new benchmark for the field.
Why it matters: WildElder provides a much-needed resource for developing and evaluating speech technologies tailored to aging populations.
A new preprint introduces the Autonomous Agency Scale (AAS), a behavioral framework designed to measure the degree of self-directed behavior in AI systems. The AAS scores systems across seven dimensions of agency, each evaluated in both active (user-initiated) and ambient (idle) temporal bands. When applied to six AI systems, the scale shows that task agents like Claude Code and Manus exhibit low ambient agency, while a persistent companion architecture uniquely demonstrates self-directed behavior during idle periods. The study also notes limitations such as single-rater assessment and potential evaluator bias.
Why it matters: The AAS provides a systematic method to distinguish between reactive and genuinely self-directed AI systems, addressing a gap in current AI evaluation frameworks.
Researchers have released the first multi-domain corpus for analyzing social biases against people experiencing homelessness (PEH), containing 1,698 gold-standard annotated texts and over 50,000 GPT-4.1-labeled texts from Reddit, X, news, and city council transcripts across ten U.S. cities (2015-2025). Benchmarking six large language models (LLMs) on this dataset revealed moderate F1 scores but significant miscalibration, such as consistent over-tagging of 'not in my backyard' (NIMBY) bias and under-detection of factual claims. The new corpus and audit protocol are intended to support municipal stigma monitoring, with caution against treating LLM-generated labels as definitive.
Why it matters: This work introduces a systematic resource and methodology for tracking and auditing social biases against a vulnerable population, potentially informing policy and public discourse.
A large-scale study analyzing 128,569 naturalistic human-LLM conversations found that informal learning behaviors, such as cognitive engagement, occurred in 31.9% of user turns, while deeper constructive engagement was present in 4.9%. The research identified that scaffolded assistant support is associated with richer, learning-oriented participation, and that these behaviors are selectively and conditionally organized. The findings suggest that human-LLM interactions can foster opportunities for users to reason and construct understanding, rather than merely serving as cognitive offloading.
Why it matters: This research highlights the potential for AI systems to support user learning and cognitive engagement, prompting a shift in evaluation metrics beyond simple answer delivery.
A survey of design students at Politecnico di Milano found very high GenAI usage, especially in the early stages of projects. The study reports that this usage does not affect students' perceptions of project ownership or creativity. Analysis of AI journals from a class showed that students have limited trust in GenAI, leading them to systematically verify and augment AI-generated outputs.
Why it matters: This study offers empirical evidence on how design students are thoughtfully integrating GenAI into their creative processes, which can inform educational strategies and tool development.
A new preprint introduces the 'Optimization Trilemma' in decentralized multi-agent coordination, focusing on the simultaneous optimization of system-wide efficiency, individual comfort, and fairness. The authors present a novel model that addresses all three objectives without significant increases in communication or computational overhead. Experiments on two real-world datasets demonstrate that the approach achieves fairer outcomes while meeting agent preferences and system goals.
Why it matters: This work advances decentralized AI by enabling fairer and more efficient resource allocation among agents without added complexity.
A preprint study in an undergraduate Probability and Statistics course compared three groups: no LLM access, unrestricted LLM access, and guided LLM access with explicit training on reasoning-focused help-seeking and stepwise hints. Students with guided LLM access demonstrated stronger independent quiz performance than those with unrestricted or no access, while unrestricted access mainly aided practice completion. The findings indicate that simply providing LLM access is insufficient for fostering independent learning; structured guidance is necessary to promote reasoning and deeper understanding.
Why it matters: This research highlights the importance of scaffolding LLM use in educational settings to enhance students' independent reasoning and learning outcomes.
A new study demonstrates that predictors from Joint-Embedding Predictive Architectures (JEPAs) can be transferred to non-JEPA encoders such as CLIP and DINOv2 using a single linear projection. This approach significantly improves classification accuracy under heavy occlusion, with the frozen JEPA predictor boosting Stanford Dogs accuracy from 15.9% to 52.1% when paired with CLIP. The benefit increases with the degree of occlusion, though the linear projection is less effective at low occlusion levels.
Why it matters: This work shows that JEPA predictors can serve as portable operators for occluded feature completion, potentially enabling more robust classification from partial views without retraining.
Researchers introduce SaaF, a novel 3D language field based on Gaussian Splatting, designed to improve interactive object retrieval in real-world scenes using natural language. SaaF addresses limitations of prior methods by employing metric learning to enhance instance discrimination and by training on multiple text labels, including ambiguous descriptions, to better handle ambiguous queries. Experiments show that SaaF achieves higher retrieval accuracy and can robustly detect and manage ambiguity in user queries.
Why it matters: This work represents a meaningful advance in enabling service robots to more accurately and interactively retrieve objects in complex environments using natural language, even when queries are ambiguous.
A preprint study based on interviews with junior and senior software engineers in South Korea suggests that generative AI is redirecting entry-level work into senior-AI workflows, potentially depriving juniors of the 'productive struggle' needed to develop expertise. The research identifies three main consequences: loss of learning opportunities for juniors, normalization of generative AI use in university classrooms, and a perceptual gap between seniors and juniors that hinders correction of these trends. The authors argue that these dynamics could undermine the traditional pathway for developing senior engineers.
Why it matters: This research raises concerns that generative AI could disrupt the established career progression in software engineering, with possible long-term impacts on the availability of experienced engineers.
A new preprint demonstrates that embeddings of image captions from language models can predict human brain activity in high-level visual regions. The study finds that machine-generated captions often outperform human-annotated ones, and that text embedders surpass autoregressive language models in both brain predictivity and alignment with human image-similarity judgments. The results also show that both the content of captions and the choice of language model significantly affect brain- and behavior-modelling performance.
Why it matters: This work highlights caption embeddings as a promising tool for probing high-level visual perception and underscores the importance of both caption content and language model architecture in modeling brain responses.
Researchers present CLARE, a clarification-aware 3D agent designed to address intent asymmetry in 3D asset creation. CLARE treats vague or underspecified user instructions as opportunities for strategic dialogue, decouples its generation pipeline into four cognitive roles, and self-evolves its clarification policy through simulated multi-turn interactions. On the new 3D-Clarify benchmark, CLARE achieves state-of-the-art success rates of 60.40% for single-step and 43.34% for multi-step tasks, more than doubling existing baselines.
Why it matters: This work significantly advances 3D asset creation by enabling agents to proactively clarify ambiguous instructions, leading to much higher task completion rates than previous approaches.
Researchers introduce Med-OPD, a post-training framework that combines on-policy distillation with medical evidence-aware supervision for medical vision-language models (Med-VLMs). The approach uses a Medical Evidence Advantage (MEA) signal to focus training on diagnosis-critical tokens and evidence-dependent reasoning. Experiments on OmniMedVQA subsets show that Med-OPD outperforms standard supervised fine-tuning and on-policy distillation methods across multiple medical imaging tasks.
Why it matters: This work offers a novel method to improve the reliability of medical vision-language models by encouraging them to base clinical reasoning on visual evidence rather than language priors.
Researchers have introduced xperception, a zero-shot 6D pose estimation system for robotic grasping that leverages CAD models and foundation models such as DINOv2 and GeDi. The system achieves millimeter-accurate pose estimation without object-specific fine-tuning or data annotation, and demonstrates robustness to occlusions in industrial tasks like bin picking. xperception is validated at TRL 6 and is engineered for deployment on industrial edge hardware, including NVIDIA Jetson Thor.
Why it matters: This approach could streamline robotic automation in flexible manufacturing by removing the need for retraining or data collection when new objects are introduced.
Researchers have introduced PriVE-Bench, a benchmark that uses paired original and counterfactual images to test whether vision-language models (VLMs) base their answers on actual visual evidence or rely on learned priors. Alongside, PriVE-Tools evaluates if providing additional tool-derived visual evidence—such as bounding boxes, crops, and contours—improves the models' grounding. The study finds that while such tools can help VLMs use visual evidence more effectively in some cases, they do not universally prevent models from defaulting to prior-based errors.
Why it matters: This work offers a systematic approach to diagnosing and addressing a key limitation in VLMs, which is essential for building more trustworthy vision-language systems.