Switch language한국어
Back to the list

The Unlearnability Phenomenon in RLVR for Language Models

TL;DR AI

Key summary

2 min read
  1. Researchers found that some hard RLVR training examples for language models remain unlearnable even when correct rollouts exist.

  2. The paper links these failures to low cross-example gradient similarity and flawed internal representations, not just weak optimization.

  3. Common fixes such as better sampling, optimizer changes, and data augmentation did not resolve the problem.

  4. The result suggests a structural limit in current RL-based reasoning training, where better training alone may not recover every hard case.

Read the original