Learning to Adapt SFT Data for Better Reasoning Generalization

TL;DR AI
2 min readKey summary
Researchers introduced DART, a method that adapts supervised fine-tuning data to better match a model’s distribution.
DART uses reinforcement learning to transform existing SFT demonstrations into model-adapted supervision before fine-tuning.
The approach improves reasoning generalization across multiple language models and datasets.
It offers a practical way to make fixed training data more effective, without relying on reinforcement learning alone.
