How to Fine-Tune a Reasoning Model? A Teacher-Student Cooperation Framework to Synthesize Student-Consistent SFT Data
TL;DR AI
2 min readKey summary
Researchers proposed TESSY, a teacher-student cooperative data-synthesis framework for reasoning models.
TESSY splits sample generation so the teacher provides capability tokens and the student provides style tokens.
This creates student-consistent SFT data and shifts training toward on-policy data instead of teacher-only off-policy data.
The approach aims to fine-tune reasoning models while preserving complex-task performance and reducing catastrophic forgetting.
