← Back to brief
ResearchOfficialPreprintarXiv Machine Learning

C-index illusion: High discrimination can mask severe calibration failures in survival models

A new arXiv preprint reproduces three published machine learning survival models from diverse domains and finds that models with high discrimination (C-index) can still exhibit severe calibration failures. In one case, a model with a C-index of 0.9595 failed a formal calibration test at an extremely significant level (p=2.6e-136). The study highlights that relying solely on discrimination metrics like the C-index can give a misleading impression of model reliability.

Why it matters: This work raises broad concerns about the evaluation of survival models in real-world applications, urging the adoption of calibration-aware metrics to avoid misplaced confidence in model predictions.

Full story at: arXiv Machine Learning