← Back to brief
ResearchOfficialPreprintarXiv AI/ML

Structured Synthetic Reasoning Data Boosts Arithmetic Accuracy in Small Language Models

Researchers fine-tuned Qwen3-0.6B and Qwen3-1.7B language models on a 21,250-example synthetic arithmetic dataset generated by GPT-5-mini. This fine-tuning improved exact-match accuracy on the GSM8K benchmark from 36.5% to 49.1% for Qwen3-0.6B and from 53.5% to 66.5% for Qwen3-1.7B. The fine-tuned Qwen3-1.7B model also showed strong transfer performance on MultiArith (98.9%) and SVAMP (73.0%), compared to much lower scores for the base model. The study demonstrates that structured synthetic reasoning data can substantially enhance arithmetic reasoning in small language models under consumer hardware constraints.

Why it matters: This work demonstrates a practical and effective approach to improving arithmetic reasoning in small language models, making them more capable for local use on consumer devices.

Full story at: arXiv AI/ML