The Flip Side of RLHF: On-Policy Feedback for Reward Model Self-Supervised Improvement
TL;DR AI
2 min readKey summary
Researchers introduced SAVE, a self-supervised framework to improve reward models for RLHF.
It scores on-policy responses with a value function, filters ambiguous cases, and trains the reward model with a contrastive objective.
SAVE reported strong results on six benchmarks and consistent gains across multiple RL algorithms and policy backbones.
The approach could reduce reliance on costly human preference labels and make reward-model training more stable as policies evolve.
