Why Do Reasoning Models Lose Coverage? The Role of Data and Forks in the Road
TL;DR AI
2 min readKey summary
A study finds that reasoning models can lose answer diversity after supervised fine-tuning.
The shrinkage is linked to training data that overemphasizes fork-in-the-road decision-point examples.
Controlled case studies support the diagnosis and show why top-1 accuracy can rise while pass@k falls.
Synthetic decision-point data and diversity-aware decoding can partially recover coverage and solution variety.
