Calibrated, abstaining coronary disease risk
A risk model in front of a clinician must have literal probabilities, and it must know when to defer. Does isotonic recalibration plus selective prediction deliver that on a small real cohort, and what does abstention buy per case deferred?
303 real patients from the Cleveland Clinic: 13 clinical features from age and chest pain type through fluoroscopy vessel count. Outcome is angiographically confirmed coronary disease, prevalence 46 percent. Six missing values, median imputed.
L2 regularized logistic regression solved by Newton's method. Isotonic calibration by pool adjacent violators, fit on a held out quarter of each training fold so it never sees test data. Abstention inside a widening band around 0.5. Bootstrap confidence intervals, 2000 resamples.
Recalibration backfired: with roughly 55 calibration points per fold, the isotonic step function overfit and calibration error doubled. The raw logistic model was already nearly calibrated. Meanwhile abstention worked exactly as designed: declining the most uncertain 43 percent of cases lifted accuracy on the answered cases from 82 to 95 percent. We kept both findings.
| AUC, 95% bootstrap CI | 0.897 · [0.860, 0.930] |
| Brier score | 0.126 |
| Calibration error, raw | 0.052 |
| Calibration error, after isotonic | 0.099 ▲ worse |
| Accuracy, full coverage | 82.2% |
| Accuracy at 57% coverage | 90.8% |
| Accuracy at 34% coverage | 95.2% |