Benchmark · general

HealthBench

16 results 6 models

OpenAI's physician-graded health-conversation benchmark.

0 14 28 42 56 2025-08-07 2026-02-08 2026-08-13 GPT-5 · 46.2 · 2025-08-07 VITA · 51.9 · 2026-08-13 VITA · 51.9 · 2026-08-13 VITA · 51.9 · 2026-08-13 GPT-5.4 · 46.1 · 2026-08-13 GPT-5.4 · 46.1 · 2026-08-13 GPT-5.4 · 46.1 · 2026-08-13 o4-mini · 44.3 · 2026-08-13 o4-mini · 44.3 · 2026-08-13 o4-mini · 44.3 · 2026-08-13 Gemini 3.1 Pro · 42.6 · 2026-08-13 Gemini 3.1 Pro · 42.6 · 2026-08-13 Gemini 3.1 Pro · 42.6 · 2026-08-13 Claude Sonnet 4.6 · 37.3 · 2026-08-13 Claude Sonnet 4.6 · 37.3 · 2026-08-13 Claude Sonnet 4.6 · 37.3 · 2026-08-13
GPT-5 VITA GPT-5.4 o4-mini Gemini 3.1 Pro Claude Sonnet 4.6
Timeline
Date Model Score Source
2026-08-13 VITA 51.9% VITA clinical RAG matches or outperforms frontier LLMs on HealthBench
2026-08-13 GPT-5.4 46.1% VITA clinical RAG matches or outperforms frontier LLMs on HealthBench
2026-08-13 o4-mini 44.3% VITA clinical RAG matches or outperforms frontier LLMs on HealthBench
2026-08-13 Gemini 3.1 Pro 42.6% VITA clinical RAG matches or outperforms frontier LLMs on HealthBench
2026-08-13 Claude Sonnet 4.6 37.3% VITA clinical RAG matches or outperforms frontier LLMs on HealthBench
2026-08-13 VITA 51.9% Corpus-specific VITA RAG matches or beats frontier LLMs on HealthBench
2026-08-13 GPT-5.4 46.1% Corpus-specific VITA RAG matches or beats frontier LLMs on HealthBench
2026-08-13 o4-mini 44.3% Corpus-specific VITA RAG matches or beats frontier LLMs on HealthBench
2026-08-13 Gemini 3.1 Pro 42.6% Corpus-specific VITA RAG matches or beats frontier LLMs on HealthBench
2026-08-13 Claude Sonnet 4.6 37.3% Corpus-specific VITA RAG matches or beats frontier LLMs on HealthBench
2026-08-13 VITA 51.9% VITA clinical RAG system matches or outperforms frontier LLMs on HealthBench
2026-08-13 GPT-5.4 46.1% VITA clinical RAG system matches or outperforms frontier LLMs on HealthBench
2026-08-13 o4-mini 44.3% VITA clinical RAG system matches or outperforms frontier LLMs on HealthBench
2026-08-13 Gemini 3.1 Pro 42.6% VITA clinical RAG system matches or outperforms frontier LLMs on HealthBench
2026-08-13 Claude Sonnet 4.6 37.3% VITA clinical RAG system matches or outperforms frontier LLMs on HealthBench
2025-08-07 GPT-5 46.2% OpenAI launches GPT-5 with adaptive reasoning and unified architecture