A study compared the clinical AI system Doctorina against eight physicians and four standalone frontier language models across 150 synthetic Polish-language primary-care consultations.

  • Doctorina achieved 82.0% Top-1 diagnostic concordance, significantly surpassing the 57.0% performance of physicians.
  • The AI system reached 97.3% primary-or-reference-differential concordance compared to 85.0% for doctors.
  • Normalized workup and treatment scores were higher for Doctorina (89.4 and 83.7) than for physicians (66.9 and 61.2).
  • Kimi K3 ranked next in diagnostics, while Claude Opus 5 led management estimates among the closely spaced top performers.

The results indicate that Doctorina's advantage extends from primary diagnosis selection to higher-rated diagnostic workup and initial treatment following adaptive consultation.