VITA, a retrieval-augmented generation system designed for low- and middle-income settings, ranks first on the English-language subset of HealthBench, outperforming GPT-5.4, o4-mini, Gemini 3.1 Pro, and Claude Sonnet 4.6.
- VITA scored 51.9% of possible rubric points on 4,023 questions, ahead of GPT-5.4 (46.1%) and Claude Sonnet 4.6 (37.3%).
- When re-evaluated against newer models like GPT-5.5 using a neutral judge, VITA achieved parity with GPT-5.5 on mean per-question score while leading on points-weighted score.
- The system retrieves from a curated corpus of disease-specific guidelines and India-specific antimicrobial resistance data.
The results indicate that purpose-built clinical RAG systems remain competitive with frontier LLMs, demonstrating that corpus specificity improves grounding despite lower communication polish.