The authors introduce Fathom-Vaidya, a 30B parameter model that uses synthetic data and rubric-based reinforcement learning to enhance diagnostic and clinical healthcare reasoning. The training framework first applies rule-guided RL to MedBullets-derived questions for diagnosis, then utilizes 5.3k synthetic multi-turn scenarios with multi-dimensional rubrics for interactive clinical tasks.

  • Achieves 50.1% accuracy on HealthBench-Hard, surpassing proprietary baselines like GPT-5 (thinking).
  • Delivers over 10% improvement on the MedXpertQA benchmark.
  • Addresses weaknesses in complex diagnostic scenarios and contextual dialogue identified by recent benchmarks.

The results demonstrate that targeted synthetic datasets and rubric-based training can systematically improve both diagnostic and interactive clinical reasoning in medical LLMs.