← Back to brief
ResearchOfficialPreprintarXiv Computation and Language

MedDDC-Eval: A Diagnosis-Decoupled Benchmark for Multi-Turn Medical Consultation Agents

Researchers introduce MedDDC-Eval, a benchmark that decouples the evaluation of information-gathering policies from diagnosis generation in multi-turn medical consultation agents. By using a shared frozen diagnostic reader, the framework enables fairer comparison of evidence acquisition strategies. The study demonstrates that changing the diagnostic model alone can shift F1 scores by up to 19 points and reverse policy rankings, highlighting the importance of decoupled evaluation. Additionally, training policies with Group Relative Policy Optimization (GRPO) on this benchmark led to notable improvements in total scores on two test splits.

Why it matters: This work provides a more rigorous and interpretable framework for evaluating and improving medical AI agents by isolating evidence collection from diagnosis, enabling clearer attribution of performance.

Full story at: arXiv Computation and Language