Researchers introduce OSReward, a realistic benchmark designed to evaluate the reliability of vision-language models (VLMs) acting as judges for computer-using agent trajectories. The study reveals that even state-of-the-art VLMs suffer from systematic leniency bias and are often too expensive for scale, while affordable open models perform poorly.
- OSReward features diverse agent trajectories with ground-truth verdicts derived from multi-stage human annotation across platforms.
- The benchmark includes OSReward-Hard for challenging cases and OSReward-Multi for fine-grained efficiency scoring.
- Evaluation shows current VLM judges mislabel failed runs as successes, creating a gap between reliability and cost.
- To address this, the authors release OS-Shepherd-100K, an open corpus of reasoning-annotated trajectory judgments.
- Trained on this data, the open OS-Shepherd models (9B and 35B) match commercial judges at 30-60% lower cost.
The work provides low-cost, stable reward signals for the community and offers extensive analyses to inform the design of reliable computer-use evaluation.