The authors introduce Fathom-Vaidya, a 30B parameter model that uses synthetic data and rubric-based reinforcement learning to enhance both diagnostic and clinical healthcare reasoning. The training framework first targets diagnostic accuracy using MedBullets-derived questions and then addresses multi-turn clinical interactions through 5.3k generated scenarios with multi-dimensional rubrics.

  • The model achieves 50.1% accuracy on HealthBench-Hard, surpassing proprietary baselines including GPT-5 (thinking).
  • It demonstrates over a 10% improvement on the MedXpertQA benchmark.
  • The approach systematically improves complex diagnostic scenarios and contextual dialogue capabilities.

This work demonstrates that targeted synthetic datasets and rubric-based training can effectively address persistent weaknesses in medical LLMs.