The authors introduce Fathom-Vaidya, a 30B parameter model that uses synthetic data and rubric-based reinforcement learning to enhance both diagnostic and clinical healthcare reasoning. The training framework first targets diagnostic accuracy using MedBullets-derived questions and then addresses multi-turn clinical interactions through 5.3k generated scenarios with multi-dimensional rubrics.
- The model achieves 50.1% accuracy on HealthBench-Hard, surpassing proprietary baselines including GPT-5 (thinking).
- It demonstrates over a 10% improvement on the MedXpertQA benchmark.
- The approach systematically improves complex diagnostic scenarios and contextual dialogue capabilities.
This work demonstrates that targeted synthetic datasets and rubric-based training can effectively address persistent weaknesses in medical LLMs.