Researchers introduce OSWorld, the first scalable, real computer environment designed for evaluating multimodal agents on open-ended tasks across Ubuntu, Windows, and macOS.

  • The benchmark includes 369 computer tasks derived from real-world use cases involving web and desktop apps, file I/O, and multi-application workflows.
  • Each task features a detailed initial state setup and custom execution-based evaluation scripts for reliable assessment.
  • Evaluations of state-of-the-art LLM/VLM-based agents show significant deficiencies, with the best model achieving only 12.24% success compared to human performance of over 72.36%.
  • The primary struggles identified are GUI grounding and operational knowledge.

This benchmark provides a unified environment for assessing agent scalability and offers insights for developing multimodal generalist agents that previous benchmarks could not support.