VITA, a retrieval-augmented generation system designed for low- and middle-income settings, ranks first on the English-language subset of HealthBench, outperforming GPT-5.4, o4-mini, Gemini 3.1 Pro, and Claude Sonnet 4.6.

  • VITA scored 51.9% of possible rubric points on 4,023 questions, ahead of GPT-5.4 (46.1%) and Claude Sonnet 4.6 (37.3%).
  • When re-evaluated against newer models like GPT-5.5 using a neutral judge, VITA achieved parity with GPT-5.5 on mean per-question score while leading on points-weighted score.
  • The system retrieves from a curated corpus of disease-specific guidelines and India-specific antimicrobial resistance data.

The results indicate that purpose-built clinical RAG systems remain competitive with frontier LLMs, demonstrating that corpus specificity improves grounding despite lower communication polish.