Not Every Rubric Teaches Equally: Policy-Aware Rubric Rewards for RLVR
TL;DR AI
2 min readKey summary
Researchers found that many rubric criteria used for RL post-training do not actually create contrastive learning signal for a given policy.
Static rubric weighting can misdirect training pressure, so the team introduced POW3R to reweight criteria during training while keeping the final evaluation target unchanged.
POW3R outperformed baseline methods across multiple policies and settings, including multimodal and HealthBench-style evaluations.
It reached similar performance with about 2.5 to 4 times fewer training steps, showing better sample efficiency.
