← Back to brief
ResearchOfficialPreprintarXiv Computation and Language

Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning

Researchers have identified a critical failure mode in long-context large language models (LLMs) called repetitive copying, where models copy input text into their reasoning traces instead of engaging in productive problem-solving. They introduce GEAR, a reward shaping method that encourages grounding in key evidence and penalizes copying from irrelevant context, leading to consistent improvements of up to +4.6 average points over standard reinforcement learning approaches across multiple benchmarks and model scales.

Why it matters: This work highlights and addresses a pervasive limitation in long-context LLMs, demonstrating that improved evidence grounding can significantly enhance reasoning performance and reduce unproductive copying.

Full story at: arXiv Computation and Language