Switch language한국어
Back to the list

Why Do Reasoning Models Lose Coverage? The Role of Data and Forks in the Road

TL;DR AI

Key summary

2 min read
  1. A study finds that reasoning models can lose answer diversity after supervised fine-tuning.

  2. The shrinkage is linked to training data that overemphasizes fork-in-the-road decision-point examples.

  3. Controlled case studies support the diagnosis and show why top-1 accuracy can rise while pass@k falls.

  4. Synthetic decision-point data and diversity-aware decoding can partially recover coverage and solution variety.

Read the original