OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models

TL;DR AI
2 min readKey summary
Researchers released OSReward, a standardized benchmark for judging computer-use agent trajectories.
The suite includes labeled data plus OSReward-Hard and OSReward-Multi subsets for tougher and more varied evaluation.
The paper finds that many vision-language model judges suffer from leniency bias, weakening evaluation reliability.
It also introduces OS-Shepherd-100K and OS-Shepherd reward models as cheaper, more dependable alternatives for scoring agents.
