Switch language한국어
Back to the list

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models

TL;DR AI

Key summary

2 min read
  1. Researchers introduced a mid-training stage where language models generate and filter multiple correct solutions before reinforcement learning.

  2. Training on diverse self-generated reasoning traces improves GRPO-based RL versus vanilla RL and STaR+RL across multiple benchmarks.

  3. Gains are larger at higher pass@k, suggesting richer priors help more when the model can explore several good paths.

  4. Analysis shows RL increasingly composes multiple heuristics over time instead of relying on a single memorized solution path.

Read the original