Switch language한국어
Back to the list

Learning to Adapt SFT Data for Better Reasoning Generalization

TL;DR AI

Key summary

2 min read
  1. Researchers introduced DART, a method that adapts supervised fine-tuning data to better match a model’s distribution.

  2. DART uses reinforcement learning to transform existing SFT demonstrations into model-adapted supervision before fine-tuning.

  3. The approach improves reasoning generalization across multiple language models and datasets.

  4. It offers a practical way to make fixed training data more effective, without relying on reinforcement learning alone.

Read the original