Reasoning Fine-Tuning Induces Persistent Latent Policy States
Researchers model Chain-of-Thought reasoning in language models as a switching dynamical system, revealing that reasoning fine-tuning globally reorganizes latent dynamics rather than merely improving token-level competence. The study finds that the recovered latent policy states show functional specialization aligned with reasoning stages, and causal interventions demonstrate their functional significance. Additionally, SDS-guided pruning of failure-prone reasoning prefixes outperforms self-consistency in 11 of 12 settings, with gains up to 12.5 percentage points.
Why it matters: This work introduces a novel mechanistic framework for analyzing and controlling reasoning in language models, with practical implications for enhancing reasoning performance through process-level interventions.
Full story at: arXiv Computation and Language ↗