Switch language한국어
Back to the list

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models

TL;DR AI

Key summary

2 min read
  1. Researchers released OSReward, a standardized benchmark for judging computer-use agent trajectories.

  2. The suite includes labeled data plus OSReward-Hard and OSReward-Multi subsets for tougher and more varied evaluation.

  3. The paper finds that many vision-language model judges suffer from leniency bias, weakening evaluation reliability.

  4. It also introduces OS-Shepherd-100K and OS-Shepherd reward models as cheaper, more dependable alternatives for scoring agents.

Read the original