Benchmark · agentic

BrowseComp

24 results 21 models

BrowseComp is OpenAI's benchmark for AI web-browsing agents, testing whether they can track down hard-to-find facts scattered across the internet. Results are reported as the percentage of questions answered correctly.

Read more
Example
A typical item asks a question whose short factual answer — a name, date, or number — is buried deep online and takes persistent, multi-step searching across many pages to locate.
Scoring
The metric is accuracy: the percentage of questions where the agent's answer matches the single known correct answer.
Verification
An answer is accepted when it matches the predetermined ground-truth answer, exact-match style; questions are built to be hard to find but easy to verify once the answer is given.
Why it matters
It probes a skill plain chatbots lack — deep, persistent web research over many steps — so it stays a demanding test of agentic browsing even for strong models.
Worked example
Task
BrowseComp item: 'Name a feature film released between 2010 and 2015 that (1) won the Academy Award for Best Picture, (2) dramatizes a single covert rescue operation, (3) was directed by the same person who plays its lead role, and (4) reaches its climax at an airport in Tehran. Give only the film's title.'
Solution
The constraints intersect at one film — a Best Picture winner from 2010–2015 that is a director-led covert-rescue drama climaxing at a Tehran airport. Final answer: Argo (2012), directed by and starring Ben Affleck.
Walkthrough
Each clue prunes the candidate set — Best Picture winners in 2010–2015, then those that are director-starred rescue dramas, then the one climaxing at Tehran's airport — leaving only Argo, which Ben Affleck both directed and starred in. BrowseComp grades a short free-text answer against the reference by exact/model-checked match, so 'Argo' scores correct.
0 24.5 49 73.5 98 2025-04-10 2025-11-27 2026-07-17 GLM-4.5 · 26.4 · 2025-08-06 Tongyi DeepResearch · 43.4 · 2025-09-16 Kimi K2 Thinking · 60.2 · 2025-11-07 MiroThinker v1.0 (72B variant) · 47.1 · 2025-11-14 Kimi K2.5 · 74.9 · 2026-01-27 Claude Opus 4.6 · 84 · 2026-02-05 Claude Opus 4.6 · 83.7 · 2026-02-05 Claude Opus 4.6 · 83.7 · 2026-03-06 Step 3.5 Flash · 69 · 2026-02-10 Qwen3.5-397B-A17B · 78.6 · 2026-02-16 Opus 4.6 · 84 · 2026-04-21 GPT-5.4 · 82.7 · 2026-04-21 GPT-5.4 Pro · 89.3 · 2026-04-21 GPT-5.5 · 84.4 · 2026-04-23 Gemini 3.1 Pro · 85.9 · 2026-04-23 Agents-A1 · 75.5 · 2026-06-30 GPT-5.6 Sol · 92.2 · 2026-07-09 Mach-Mind-4-Flash · 72.3 · 2026-07-13 Mach-Mind-4-Flash · 72.3 · 2026-07-13 Qwen3-4B · 35.6 · 2026-07-16 Qwen3-30B-A3B · 42.6 · 2026-07-16 GPT-Live-1 · 75.2 · 2026-07-17 Kimi K3 · 91.2 · 2026-07-17 OpenAI Deep Research · 51.5 · 2025-04-10
GLM-4.5 Tongyi DeepResearch Kimi K2 Thinking MiroThinker v1.0 (72B variant) Kimi K2.5 Claude Opus 4.6 Step 3.5 Flash Qwen3.5-397B-A17B Opus 4.6 GPT-5.4 GPT-5.4 Pro GPT-5.5 Gemini 3.1 Pro Agents-A1 GPT-5.6 Sol Mach-Mind-4-Flash Qwen3-4B Qwen3-30B-A3B GPT-Live-1 Kimi K3 OpenAI Deep Research
Timeline
Date Model Score Source
2026-07-17 Kimi K3 91.2% Moonshot AI releases open Kimi K3, a 2.8T-parameter MoE model with 1M context
2026-07-17 GPT-Live-1 75.2% OpenAI releases full-duplex voice models and German court holds Google liable for AI Overviews
2026-07-16 Qwen3-4B 35.6% TRACE improves long-horizon agent tool-use via turn-level credit estimation
2026-07-16 Qwen3-30B-A3B 42.6% TRACE improves long-horizon agent tool-use via turn-level credit estimation
2026-07-13 Mach-Mind-4-Flash 72.31% Mach-Mind-4-Flash: 35B MoE agentic model matches larger models via post-training optimization
2026-07-13 Mach-Mind-4-Flash 72.31% Mach-Mind-4-Flash: 35B MoE agentic model matches larger models via post-training optimization
2026-07-09 GPT-5.6 Sol 92.2% OpenAI launches GPT-5.6 family with Sol, Terra, and Luna models
2026-06-30 Agents-A1 75.5pts Scaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B Agent
2026-04-23 GPT-5.5 84.4% OpenAI releases GPT-5.5 with agentic capabilities and doubled API pricing
2026-04-23 Gemini 3.1 Pro 85.9% OpenAI releases GPT-5.5 with agentic capabilities and doubled API pricing
2026-04-21 Opus 4.6 84.0% Google launches Deep Research and Deep Research Max agents built on Gemini 3.1 Pro
2026-04-21 GPT-5.4 82.7% Google launches Deep Research and Deep Research Max agents built on Gemini 3.1 Pro
2026-04-21 GPT-5.4 Pro 89.3% Google launches Deep Research and Deep Research Max agents built on Gemini 3.1 Pro
2026-03-06 Claude Opus 4.6 83.73% Anthropic releases Claude Opus 4.6 system card detailing capabilities and safety assessments
2026-02-16 Qwen3.5-397B-A17B 78.6% Qwen releases FP8-quantized Qwen3.5-397B-A17B model
2026-02-10 Step 3.5 Flash 69.0% Step 3.5 Flash: open 11B-active MoE model achieves frontier-level agentic performance
2026-02-05 Claude Opus 4.6 84.0% Anthropic releases Claude Opus 4.6 with 1M context window and agentic improvements
2026-02-05 Claude Opus 4.6 83.73% Anthropic publishes Claude Opus 4.6 system card detailing capabilities and safety evaluations
2026-01-27 Kimi K2.5 74.9% Moonshot releases Kimi K2.5 with Agent Swarm technology
2025-11-14 MiroThinker v1.0 (72B variant) 47.1% MiroThinker v1.0 introduces interactive scaling for open-source research agents
2025-11-07 Kimi K2 Thinking 60.2% Moonshot AI releases Kimi K2 Thinking, an open-source 1T parameter reasoning model
2025-09-16 Tongyi DeepResearch 43.4% Tongyi releases DeepResearch, an open-source web agent matching proprietary performance
2025-08-06 GLM-4.5 26.4% Zhipu AI releases GLM-4.5 models matching Claude and DeepSeek
2025-04-10 OpenAI Deep Research 51.5% OpenAI