Researchers introduce ClinFusion, a vision-centric multimodal large language model designed for holistic medical understanding that unifies 2D and native 3D image analysis. The system utilizes a compositional cascaded vision encoder with a Cascade Spatial-Aware Locality Fusion operator to address the challenges of heterogeneous medical imaging.
- ClinFusion outperforms leading open-source models like Hulu-Med and Lingshu on 20 out of 24 benchmarks, including visual question answering and report generation.
- It demonstrates multimodal capabilities superior to proprietary models such as GPT-5.2 and Gemini-3-Flash on 13 out of 16 benchmarks.
- The authors propose MedIF-Bench for instruction-following assessment and a region-of-interest-grounded method for factualness-driven report evaluation.
- A blinded evaluation by board-certified radiologists confirmed ClinFusion produced the highest-ranked reports, validating the new RoI-grounded metric's correlation with expert judgment.
ClinFusion establishes a new state of the art across comprehensive medical benchmarks and can be augmented with agentic tool use for retrieval-augmented clinical workflows.