Switch language한국어
Back to the list

How to Fine-Tune a Reasoning Model? A Teacher-Student Cooperation Framework to Synthesize Student-Consistent SFT Data

TL;DR AI

Key summary

2 min read
  1. Researchers proposed TESSY, a teacher-student cooperative data-synthesis framework for reasoning models.

  2. TESSY splits sample generation so the teacher provides capability tokens and the student provides style tokens.

  3. This creates student-consistent SFT data and shifts training toward on-policy data instead of teacher-only off-policy data.

  4. The approach aims to fine-tune reasoning models while preserving complex-task performance and reducing catastrophic forgetting.

Read the original