← Back to brief
ResearchOfficialPreprintarXiv Machine Learning

Off-Policy Distillation Study Reveals Capability Trade-offs and Adaptive Objective Routing in LLM Pre-training

A new preprint analyzes off-policy distillation in large language model (LLM) pre-training by decomposing the process into objective-to-capability and data-to-objective analyses. The study finds that language-modeling and knowledge-distillation objectives lead to different capability profiles, and introduces diagnostic metrics to quantify these differences. Notably, domain-level adaptive objective routing—applying different objectives to different data domains—outperforms single-objective baselines, while token-level routing does not. The work reframes continued pre-training as a structured supervision-design problem.

Why it matters: This research offers a systematic framework for understanding and optimizing the interaction between training objectives and data in LLM pre-training, with practical implications for improving model capabilities.

Full story at: arXiv Machine Learning