Researchers introduce OSReward, a realistic benchmark designed to evaluate the reliability of vision-language models (VLMs) acting as judges for computer-using agent trajectories. The study reveals that even state-of-the-art VLMs suffer from systematic leniency bias and are often too expensive for scale, while affordable open models perform poorly.

  • OSReward features diverse agent trajectories with ground-truth verdicts derived from multi-stage human annotation across platforms.
  • The benchmark includes OSReward-Hard for challenging cases and OSReward-Multi for fine-grained efficiency scoring.
  • Evaluation shows current VLM judges mislabel failed runs as successes, creating a gap between reliability and cost.
  • To address this, the authors release OS-Shepherd-100K, an open corpus of reasoning-annotated trajectory judgments.
  • Trained on this data, the open OS-Shepherd models (9B and 35B) match commercial judges at 30-60% lower cost.

The work provides low-cost, stable reward signals for the community and offers extensive analyses to inform the design of reliable computer-use evaluation.