Lost in the Folds: When Cross-Validation Is Not a Deep Ensemble for Uncertainty Estimation
TL;DR AI
2 min readKey summary
A review of recent uncertainty-estimation papers in medical image segmentation highlights a key mismatch in how ensemble methods are used and interpreted.
Comparing a 5-fold cross-validation ensemble with a standard deep ensemble, researchers found the deep ensemble performed better for calibration and failure detection.
The cross-validation ensemble, however, more closely reflected rater disagreement and inter-observer ambiguity in multi-rater datasets.
The findings clarify that these two ensemble types support different uncertainty goals, which matters for reliability claims in clinical AI systems.
