CADENCE: Closing the Reasoning Gap via Coverage-Adaptive On-Policy Distillation
TL;DR AI
2 min readKey summary
Researchers introduced CADENCE, a unified on-policy distillation framework for transferring reasoning into small language models.
It combines adaptive KL scheduling with auxiliary mechanisms to reduce cold-start failures, improve scheduling, and address sparse rewards.
On GSM8K and MATH-500, 0.5B students outperformed prior matched-compute baselines, using only a single Mac Studio.
The work suggests a more stable, compute-efficient path for distilling math reasoning from large models into much smaller ones.
