A purpose-built retrieval-augmented generation (RAG) system named VITA, designed for low- and middle-income settings, was evaluated against general-purpose large language models on the HealthBench benchmark. On 4,023 questions scored by a GPT-4.1 judge, VITA ranked first with 51.9% of possible rubric points, surpassing GPT-5.4, o4-mini, Gemini 3.1 Pro, and Claude Sonnet 4.6.
When re-evaluated against newer models using a neutral open-weight judge, VITA achieved parity with GPT-5.5 on mean per-question score while leading on points-weighted score and winning the most questions. The study indicates that corpus specificity improves grounding in clinical contexts, allowing specialized systems to remain competitive with frontier LLMs despite potentially lower communication polish.