Researchers introduce ClinFusion, a vision-centric multimodal large language model designed to address the challenges of deploying AI in clinical practice by unifying 2D and 3D medical image understanding. The system features a compositional cascaded vision encoder with a Cascade Spatial-Aware Locality Fusion operator and includes a new evaluation framework comprising MedIF-Bench and region-of-interest-grounded metrics.

  • ClinFusion sets a new state-of-the-art across 20 out of 24 benchmarks, outperforming open-source models like Hulu-Med and Lingshu.
  • It demonstrates multimodal capabilities superior to proprietary models GPT-5.2 and Gemini-3-Flash on 13 out of 16 benchmarks.
  • The model supports agentic tool use for retrieval-augmented and tool-assisted clinical workflows.
  • Blinded evaluations by board-certified radiologists confirm ClinFusion produces the highest-ranked reports.
  • The proposed RoI-grounded metric shows the strongest correlation with expert judgment among automatic evaluation metrics.

The authors consider this significant because it provides a clinically aligned, factualness-driven assessment method and validates that their model's output aligns closely with expert radiologist judgment.