Unsupervised Process Reward Models
TL;DR AI
2 min readKey summary
Researchers introduced uPRM, an unsupervised process reward model that identifies likely first error steps using next-token probabilities.
The method removes the need for human step-level labels while still supporting reasoning verification and error detection.
On ProcessBench, uPRM shows competitive or better performance than supervised and baseline approaches.
The approach also improves reinforcement learning, offering a cheaper path to training reward models for reasoning tasks.
