ACL 2026: Alibaba DAMO Academy's I2B-LPO Breaks RLVR Homogenization — From Repetitive Sampling to Effective Exploration

TL;DR AI
2 min readKey summary
Alibaba DAMO Academy introduced I2B-LPO, an exploration-enhancement framework for RLVR post-training.
Instead of relying on repetitive sampling, it encourages more diverse and informative reasoning trajectories.
The approach reportedly improves math benchmark accuracy by up to 5.3% and semantic diversity by up to 7.4%.
It highlights a key RLVR limitation: more samples do not automatically lead to better learning.
The work, from Alibaba’s Intelligent Decision team, was accepted to ACL 2026 Main.



