H2O.ai announces that its h2oGPTe Agent has achieved the #1 position on the General AI Assistants (GAIA) benchmark, reaching a 75% accuracy rate. This result marks the first time a grade of C has been attained on the GAIA test set and places the agent ahead of OpenAI's Deep Research while matching the performance of Manus.

  • The agent utilizes Anthropic's Claude 3.7 Sonnet with specialized reasoning mode for complex analytical tasks.
  • New capabilities include sophisticated browser navigation, unified search across multiple engines like Google and Bing, and GitHub integration for software engineering.
  • H2O.ai emphasizes that its results are based on the GAIA test set rather than the validation set, citing potential data contamination in the latter.
  • The platform now features live source attribution and inline response citations to enhance transparency.

The achievement validates H2O.ai's approach to creating AI assistants capable of handling real-world tasks requiring reasoning, multi-modality, and tool use.