← Back to brief
ResearchOfficialPreprintarXiv Computation and Language

ClinMM-Bench: Large-Scale Benchmark Exposes Gaps in Multimodal LLM Clinical Reasoning

A new arXiv preprint introduces ClinMM-Bench, a large-scale benchmark for evaluating multi-turn, multimodal clinical diagnostic reasoning in AI models. The benchmark includes over 1,000 real-world cases and thousands of medical images, testing 15 multimodal large language models (MLLMs). While proprietary models performed best, all models struggled to consistently provide fully correct diagnoses and reliable reasoning, with error analysis revealing several recurring failure modes.

Why it matters: This work highlights significant limitations in current multimodal LLMs for complex clinical reasoning, underscoring challenges for safe AI use in healthcare.

Full story at: arXiv Computation and Language