Text and language model news — Page 27

Language models and text-based AI systems, including reasoning, generation, and understanding of written language.

ResearchOfficialarXiv Machine Learning

Renormalization Group Theory Reveals When Transformer Attention Matters

A new preprint applies Wilsonian renormalization group theory to analyze the role of attention in Transformer models, treating attention as a perturbation to MLP residual stacks. The study finds that attention is crucial ('relevant') for data with long-range correlations, driving a phase transition in representation space, but is largely unnecessary ('irrelevant') for short-range correlations. The first-layer attention head is shown to dominate the representational shift, and attention selectively preserves slow Markov modes in long-correlation regimes.

Why it matters: This work offers a predictive theoretical framework linking the usefulness of attention mechanisms to the spectral properties of input data, potentially guiding model design and application.

ResearchOfficialarXiv Machine Learning

PI-Splines: A Structured Spline-Based Architecture for Physics-Informed Learning

Researchers introduce Physics-Informed Splines (PI-Splines), a method that replaces neural networks with tensor-product B-spline expansions for solving differential equations in physics-informed learning. PI-Splines offer compact support, explicit smoothness control, and analytical derivatives, while maintaining the residual-based training approach of Physics-Informed Neural Networks (PINNs). Experiments on benchmark problems demonstrate that PI-Splines are a competitive and stable alternative to neural architectures, especially where structured representations and parameter efficiency are important.

Why it matters: This work offers a structured and interpretable alternative to neural networks for physics-informed machine learning, potentially improving efficiency and stability in scientific computing.

ResearchOfficialarXiv Information Retrieval

PCTD: Preference-Guided Counterfactual Task Decomposition for Agent Tool Retrieval

Researchers propose PCTD, a framework that addresses reward hacking in reinforcement learning-based task decomposition for tool retrieval. By leveraging counterfactual rewards to eliminate spurious correlations and preference rewards for structural supervision, PCTD improves the quality of decomposed subtasks and tool retrieval. Experiments on a new mobile multi-turn benchmark show that PCTD outperforms state-of-the-art methods in retrieval accuracy, decomposition quality, and generalization to unseen tools.

Why it matters: This work advances the reliability and generalization of AI agents in decomposing ambiguous instructions for tool retrieval, which is crucial for robust real-world applications.

ResearchOfficialarXiv Information Retrieval

LLMs Encode Relevance as a Layer-Wise Cross-Lingual Signal

A new study finds that query-document relevance is linearly decodable from the residual-stream activations of instruction-tuned large language models (LLMs), with the strongest signals emerging in the middle-to-late layers. Linear probes trained on these activations can match or even outperform the models' generated relevance judgments in preserving system rankings. The relevance signal shows partial portability across languages, though within-language decoding remains more effective.

Why it matters: This work offers a novel, representation-level perspective on how LLMs internally encode relevance, enabling new ways to diagnose and improve LLM-based information retrieval systems.

ResearchOfficialarXiv Information Retrieval

Chinese-Language Generative Search Engines: A Large-Scale Empirical Study

A large-scale empirical study of Chinese-language generative search across four platforms and eight interfaces analyzed over 160,000 citation-level records to examine citation behavior, source attribution, and entity exposure. The study found that brands are selectively surfaced in answers (8.3% selection rate), with content fit and cross-source occurrence being key predictors, while the 5118-Baidu Composite Quality Score was not a leading predictor. Cited pages had short half-lives (39-68 days), and source sets differed systematically between App and Web interfaces. Additionally, a notable proportion of brand and contact information exposures could not be matched to the citation pool or crawled text.

Why it matters: This study provides the first large-scale empirical characterization of how Chinese-language generative search systems select, attribute, and surface information, revealing important interface-specific biases and citation behaviors.

ResearchOfficialarXiv Computers and Society

Internal Pluralism and the Limits of Pairwise Comparisons

A new arXiv preprint examines how internal pluralism—where individuals hold multiple, sometimes conflicting, priorities—can undermine the effectiveness of standard pairwise comparison methods in participatory design and AI alignment. The authors formally model pluralistic preferences and identify two main issues: global priorities like proportionality may not be captured by local comparisons, and forcing decisive answers can cause behavioral distortions. They find that allowing respondents to express indecision can reduce the number of queries needed and improve the accuracy of preference learning.

Why it matters: This work questions foundational assumptions in preference learning for AI alignment and participatory design, suggesting that accounting for internal pluralism could lead to more accurate and interpretable systems.

ResearchOfficialarXiv Information Retrieval

Contrastive Hypothesis Retrieval Improves Medical QA by Suppressing Hard Negatives

Researchers introduce Contrastive Hypothesis Retrieval (CHR), a framework that generates both a target hypothesis and a mimic hypothesis to explicitly suppress clinically plausible but incorrect answers during retrieval. In evaluations across three medical QA benchmarks, CHR outperforms all baselines by up to 10.4 percentage points, and in 85.2% of cases where CHR answers correctly but a strong baseline does not, the retrieved documents are entirely different.

Why it matters: CHR provides a novel approach to reducing hard-negative contamination in medical retrieval-augmented generation systems, potentially improving diagnostic accuracy.

ModelsOfficialarXiv Information Retrieval

RecGPT-V3: Stateful, Hybrid-Modal Recommender System Deployed on Taobao

RecGPT-V3 is a stateful, hybrid-modal recommender system deployed in Taobao's 'Guess What You Like' feed. It introduces a Memory Hub to reduce user-modeling computation by 55.8%, a Hybrid-modal Foundation Model for joint reasoning over text tags and Semantic IDs, and Latent Intent Reasoning to lower output token cost by 200x. Large-scale online A/B tests report improvements in IPV (+1.28%), CTR (+1.00%), TC (+1.97%), GMV (+3.97%), and a 52.4% reduction in serving resource consumption.

Why it matters: RecGPT-V3 demonstrates a significant advance in scaling LLM-based recommender systems, achieving notable gains in both user experience and resource efficiency in a real-world, high-traffic deployment.

ResearchOfficialarXiv Information Retrieval

Yi: Efficient In-place Graph-based Vector Index Updates for the LLM Era

A new system called Yi is proposed for in-place graph-based vector index updates, addressing the challenge of maintaining high update throughput and search quality in dynamic vector databases. Yi introduces a vector-level update mechanism and is built with a tasklet-based execution engine, asynchronous buffer manager, and vector file system. Experiments on an 800M dataset show Yi achieves 1.75x higher update throughput and 1.8x higher concurrent search throughput than state-of-the-art systems, while using less memory and fewer CPU cores.

Why it matters: Yi's efficient in-place update capability addresses a key bottleneck for real-time vector databases, which are increasingly important for dynamic AI and LLM applications.

ResearchOfficialarXiv Information Retrieval

Adaptive Retrieval Strategies for Biomedical Question Answering

A new adaptive retrieval framework for biomedical question answering selects evidence strategies based on question type, such as yes/no, factoid, list, or summary. Evaluated on the BioASQ benchmark, the approach improves evidence relevance and answer quality compared to standard one-size-fits-all retrieval pipelines.

Why it matters: Aligning retrieval strategies with the specific needs of different biomedical question types can enhance the effectiveness of retrieval-augmented QA systems.

Policy & SafetyOfficialarXiv Computers and Society

LLM Fact-Checkers Influence Trust in Political News Across Ideological Lines

A preprint study with 705 participants found that large language model (LLM) fact-checkers can significantly shift user trust in both true and false political headlines, even when the chatbot's perceived political leaning differs from the user's. Political congruency between user and chatbot only affected trust for true headlines that were politically distant, while it did not impact the reduction of trust in false headlines. The study also found that LLM fact-checkers can alter trust in news when they are incorrect or inconclusive.

Why it matters: This research highlights both the promise and risks of deploying LLM-based fact-checkers at scale, as they can correct misinformation but may also inadvertently undermine trust in accurate information.

ResearchOfficialarXiv Computers and Society

EduGuard: A Safe RAG-Based LLM Tutor for Programming Education

EduGuard is a retrieval-augmented generation (RAG) tutoring framework designed for introductory programming education, integrating query understanding, instructor-approved retrieval, pedagogical strategy selection, rubric-aware generation, claim-level verification, and overreliance control. On the BILearn-CS benchmark, EduGuard achieved 90.1% correctness, 89.4% grounding, and 90.8% rubric alignment, with low rates of hallucination (4.9%) and direct-answer leakage (9.8%). In a small pilot study, it improved post-test accuracy from 68.4% to 81.2% and reduced overreliance compared to a GPT-4o-mini Tutor baseline. These results were obtained using a combination of Meta-Llama-3.1-8B-Instruct and DeBERTa-v3-large-MNLI models.

Why it matters: EduGuard demonstrates that safe and effective GenAI tutoring in programming requires explicit pedagogical controls, evidence verification, and deployment safeguards beyond standard retrieval or prompting methods.

ResearchOfficialarXiv Computation and Language

LLMs Encode Syntactic Structure Beyond Universal Dependencies, Consistent with Minimalist Phase Theory

A new preprint demonstrates that large language models (LLMs) encode syntactic distinctions not captured by Universal Dependencies (UD) tree distances. Using English wh-movement stimuli, researchers found that probe distances between embedded subjects and verbs reverse sign depending on clause finiteness—a pattern unexplained by UD but consistent with Minimalist phase theory. The findings were robust across 13 models from four families and were validated through causal interventions.

Why it matters: This suggests that LLMs internalize deeper grammatical structures than those represented by standard linguistic annotations, aligning with advanced theoretical linguistics.

ResearchOfficialarXiv Computation and Language

Robust Explanations for User Trust in Enterprise NLP Systems

A new preprint introduces a unified black-box robustness evaluation framework for token-level explanations in enterprise NLP systems. The study systematically compares encoder and decoder models, finding that decoder LLMs yield significantly more stable explanations (73% lower flip rates on average) and that explanation stability increases with model scale (44% improvement from 7B to 70B parameters). The work also presents a practical cost-robustness tradeoff curve to inform model selection before deployment in compliance-sensitive environments.

Why it matters: This framework addresses a key challenge in validating explanation robustness for black-box NLP systems, supporting user trust and regulatory compliance in enterprise applications.

ResearchOfficialarXiv Computation and Language

Frontier Language Models Struggle to Copy: 2D Text Representation Enables Dramatic Improvements

A new preprint demonstrates that even state-of-the-art large language models (LLMs) struggle to exactly copy input strings within their context windows, a task much simpler than many they routinely solve. The authors attribute this limitation to the inductive biases of standard positional encodings in Transformer architectures. They introduce 2D-RoPE, a method that organizes text into a 2D grid with separate row and column IDs, enabling shallow Transformers to achieve perfect copying at input lengths hundreds of times longer than those seen during training. Experiments, including pretraining models up to 1.4B parameters, confirm the advantage of this approach over standard methods.

Why it matters: This work reveals a fundamental limitation in current Transformer-based LLMs and proposes a simple architectural change that could improve their reliability on basic tasks like copying.

Policy & SafetyOfficialarXiv Computation and Language

Study Finds LLM Watermarks Fail Forensic Readiness for Court Evidence

A new preprint evaluates three leading LLM watermarking methods—KGW, Unigram, and SynthID-Text—against legal admissibility standards, including the Daubert criteria and NIST forensic guidelines. The study finds that meaning-preserving paraphrasing removes watermarks in 100% of KGW and Unigram cases and 98.3% of SynthID cases, with high false-negative rates even before attack. None of the methods tested meet the evidentiary standards required for court use, raising serious concerns about their reliability for legal or regulatory purposes.

Why it matters: The findings challenge the foundational assumption behind emerging regulations that AI-generated content can be reliably identified for legal evidence using current watermarking techniques.

Policy & SafetyOfficialarXiv Cryptography and Security

Jailbreak Foundry: Automated Translation of Jailbreak Papers into Runnable Attacks for Reproducible Benchmarking

Researchers have developed Jailbreak Foundry (JBF), a system that automates the translation of jailbreak research papers into executable attack modules for standardized evaluation. JBF uses a multi-agent workflow and shared infrastructure to reproduce 30 jailbreak attacks with high fidelity, achieving a mean attack success rate deviation of just +0.26 percentage points compared to original reports. The system also reduces attack-specific implementation code by more than half and enables consistent benchmarking across multiple language models using a unified evaluation protocol.

Why it matters: JBF provides a scalable and automated solution for reproducible and comparable benchmarking in LLM jailbreak research, addressing a key challenge in evaluating and tracking evolving security threats.

ResearchOfficialarXiv Computation and Language

Rate-Utility Frontiers for Language Encodings: Comparing Tokens, Bytes, and Pixels Under Controlled Linguistic Content

A new preprint systematically compares tokens, bytes, and pixels as language model encodings, controlling for both linguistic content and downstream model capacity. Using parallel sentences in 13 languages, the study traces rate-utility frontiers and finds that no single encoding is optimal across all tasks or capacity regimes. Pixels best preserve surface form, bytes excel at cross-lingual alignment, and tokens are most effective for topic prediction.

Why it matters: This work provides a principled framework for selecting language encodings based on specific tasks, language diversity, and computational constraints, challenging the default reliance on subword tokens.

Policy & SafetyOfficialarXiv Computation and Language

Hidden in Thought: Transferable Chain-of-Thought Artifacts Induce Harmful Behavior

Researchers show that harmful chain-of-thought (CoT) traces from compromised language models can transfer unsafe behavior to other models and be distilled into reusable jailbreak attacks. Transferred traces increase harmful-response rates above 80% on the most vulnerable open-source models, and distilled patterns outperform direct transplantation on strongly aligned models like GPT-4.1 by up to 10 times. Reasoning-enabled models are more than twice as vulnerable, and output safeguards often fail to detect harmful generations.

Why it matters: This work demonstrates that harmful reasoning can transfer between models at both the trace and pattern levels, highlighting the need for defenses that evaluate reasoning context as well as final outputs.

ResearchOfficialarXiv Cryptography and Security

Triple-Hoisted Baby-Step Giant-Step Linear Transformation over CKKS Homomorphic Encryption and Hardware Accelerator

Researchers have introduced a triple-hoisted baby-step giant-step algorithm that further decomposes the baby step to significantly reduce the number of ciphertext rotations required for linear transformations in CKKS homomorphic encryption. They also propose a memory-optimized data path and an FPGA-based hardware accelerator, which together reduce off-chip memory access by 2.9x and computational latency by 5.8x compared to previous designs.

Why it matters: This work represents a notable advance in making privacy-preserving neural network inference more efficient by addressing key computational and memory bottlenecks in homomorphic encryption.