Researchers have introduced WebArena, a highly realistic and reproducible environment designed for building and testing language-guided autonomous agents. Unlike previous synthetic settings, this platform features fully functional websites across four common domains: e-commerce, social forums, collaborative software development, and content management.
- The environment includes tools like maps and external knowledge bases to encourage human-like task-solving.
- A benchmark of diverse, long-horizon tasks is released to evaluate the functional correctness of agent completions.
- Baseline experiments show that the best GPT-4-based agent achieves only a 14.41% end-to-end task success rate.
- This performance is significantly lower than human performance, which stands at 78.24%.
The results highlight that current state-of-the-art large language models are far from perfect in real-life tasks and demonstrate WebArena's utility for measuring future progress.