The authors introduce Fathom-Vaidya, a 30B parameter model that uses synthetic data and rubric-based reinforcement learning to enhance both diagnostic and clinical healthcare reasoning.
- The framework employs rule- and rubric-guided RL on MedBullets-derived questions to improve diagnostic accuracy.
- It generates 5.3k synthetic multi-turn scenarios with multi-dimensional rubrics to assess interactive clinical dialogue.
- The model achieves over 10% improvement on MedXpertQA and 50.1% accuracy on HealthBench-Hard, surpassing proprietary baselines like GPT-5 (thinking).
This approach demonstrates that targeted synthetic datasets and rubric-based training can systematically improve medical LLM performance across complex diagnostic and conversational tasks.