Selective classification under imbalance in multiclass settings: A novel metric for bias-aware risk–coverage evaluation


Sağlam F., Özgen Ü., Uygun A., Dinçer O. S., Albayrak C.

JOURNAL OF BIOMEDICAL INFORMATICS, cilt.181, ss.1-17, 2026 (SCI-Expanded, Scopus)

  • Yayın Türü: Makale / Tam Makale
  • Cilt numarası: 181
  • Basım Tarihi: 2026
  • Doi Numarası: 10.1016/j.jbi.2026.105084
  • Dergi Adı: JOURNAL OF BIOMEDICAL INFORMATICS
  • Derginin Tarandığı İndeksler: Applied Science & Technology Source, Scopus, Science Citation Index Expanded (SCI-EXPANDED), BIOSIS, CINAHL, Compendex, EMBASE, INSPEC, MEDLINE
  • Sayfa Sayıları: ss.1-17
  • Ondokuz Mayıs Üniversitesi Adresli: Evet

Özet

Objective Selective classification improves reliability by allowing models to abstain on uncertain inputs, which is critical in safety-sensitive domains such as healthcare. However, commonly used evaluation metrics obscure fairness issues under class imbalance, leading to disproportionately high rejection of minority classes and misleadingly favorable assessments of selective strategies. Methods We introduce two imbalance-aware evaluation metrics, Class-Averaged AURC (CA-AURC) and ClassAveraged AUGRC (CA-AUGRC), which integrate risk against class-specific coverage and then average across classes, alongside the Area under the IAM-coverage curve (AUIC) as a complementary imbalance-aware performance metric. In addition, we propose a class-conditional coverage-matching selection strategy that enforces balanced rejection across diagnostic categories. The proposed framework is evaluated on a clinical complete blood count dataset comprising 3316 patient records, 11 laboratory features, and 9 diagnostic classes, and validated on three additional publicly available benchmark datasets covering different domains, class structures, and imbalance levels. Results While conventional metrics such as AURC and AUGRC favor classical selective strategies, the proposed class-averaged metrics reveal substantial disparities in rejection behavior under class imbalance. Using CA-AUGRC and AUIC, the class-conditional strategy consistently outperforms both the classical approach and LABEL, a theoretically principled set-valued baseline, across all four datasets and five uncertainty measures. Rank-based statistical comparisons confirm significant advantages of the class-conditional strategy on CA-AUGRC (𝑝 = 0.009, 𝑟 = 0.638) and AUIC (𝑝 < 0.001, 𝑟 = 0.957). Conclusion The results demonstrate that both evaluation and selection in selective classification must be class-aware to ensure fairness and clinical usefulness. Class-averaged metrics and class-conditional selection provide a more reliable basis for assessing selective classifiers in imbalanced medical data, with consistent generalizability across diverse datasets and uncertainty measures.