The Unlearnability Phenomenon in RLVR for Language Models
TL;DR AI
2 min readKey summary
Researchers found that some hard RLVR training examples for language models remain unlearnable even when correct rollouts exist.
The paper links these failures to low cross-example gradient similarity and flawed internal representations, not just weak optimization.
Common fixes such as better sampling, optimizer changes, and data augmentation did not resolve the problem.
The result suggests a structural limit in current RL-based reasoning training, where better training alone may not recover every hard case.
