Call for paper last date Oct. 20, 2026, 11:59 p.m. Submit
Research Article

Confidence or Competence: A Calibration Study of Machine Learning Classifiers for Selective Prediction in High-Stakes Contexts

by  Saumyya Dalal
journal cover
International Journal of Computer Applications
Foundation of Computer Science (FCS), NY, USA
Volume 187 - Issue 143
Published: September 2026
Authors: Saumyya Dalal
10.5120/ijca0c5a8ede52b8
PDF

Saumyya Dalal . Confidence or Competence: A Calibration Study of Machine Learning Classifiers for Selective Prediction in High-Stakes Contexts. International Journal of Computer Applications. 187, 143 (September 2026), 16-22. DOI=10.5120/ijca0c5a8ede52b8

                        @article{ 10.5120/ijca0c5a8ede52b8,
                        author  = { Saumyya Dalal },
                        title   = { Confidence or Competence: A Calibration Study of Machine Learning Classifiers for Selective Prediction in High-Stakes Contexts },
                        journal = { International Journal of Computer Applications },
                        year    = { 2026 },
                        volume  = { 187 },
                        number  = { 143 },
                        pages   = { 16-22 },
                        doi     = { 10.5120/ijca0c5a8ede52b8 },
                        publisher = { Foundation of Computer Science (FCS), NY, USA }
                        }
                        %0 Journal Article
                        %D 2026
                        %A Saumyya Dalal
                        %T Confidence or Competence: A Calibration Study of Machine Learning Classifiers for Selective Prediction in High-Stakes Contexts%T 
                        %J International Journal of Computer Applications
                        %V 187
                        %N 143
                        %P 16-22
                        %R 10.5120/ijca0c5a8ede52b8
                        %I Foundation of Computer Science (FCS), NY, USA
Abstract

The central claim of the paper is that, besides being capable of making correct predictions, a model must also have the ability to say "I do not know" in critical moments. To accomplish this, the indicator of model honesty should be given more significance than the performance measure modeled in terms of accuracy. Five classifiers have been trained on the COMPAS recidivism dataset, and their performance has been evaluated not only based on accuracy and F1 but also on their calibration quality and performance of selective predictions. Results indicate that the Multilayer Perceptron (MLP), despite its modest raw accuracy, achieves the best calibration. The Random Forest model, which is often the default recommendation, is the least well-calibrated regarding its uncertainty in the uncalibrated state but improves significantly under post-hoc calibration. Every single classifier improves accuracy when allowed to abstain on low-confidence cases. The conclusion reached in the paper is that selective prediction is a practical, model-independent approach that can lead to safer AI systems, along with accuracy evaluation not being enough for making deployment decisions.

References
  • J. Angwin, J. Larson, L. Kirchner, S. Mattu, and D. Phiffer, “Machine Bias: There’s software used across the country to predict future criminals. And it’s biased against blacks,” ProPublica, May 2016. [Online]. Available: https://www.propublica.org/article/machine-bias-risk-assessments-in-criminal-sentencing.
  • C. Guo, G. Pleiss, Y. Sun, and K. Q. Weinberger, “On Calibration of Modern Neural Networks,” in Proc. 34th Int. Conf. Machine Learning (ICML), 2017, pp. 1321–1330.
  • Y. Geifman and R. El-Yaniv, “Selective Classification for Deep Neural Networks,” in Advances in Neural Information Processing Systems (NeurIPS), vol. 30, 2017, pp. 4878–4887.
  • J. Platt, “Probabilistic Outputs for Support Vector Machines and Comparisons to Regularized Likelihood Methods,” in Advances in Large Margin Classifiers, A. Smola, P. Bartlett, B. Schölkopf, and D. Schuurmans, Eds. Cambridge, MA, USA: MIT Press, 1999, pp. 61–74.
  • A. Niculescu-Mizil and R. Caruana, “Predicting Good Probabilities with Supervised Learning,” in Proc. 22nd Int. Conf. Machine Learning (ICML), 2005, pp. 625–632.
  • G. W. Brier, “Verification of Forecasts Expressed in Terms of Probability,” Monthly Weather Review, vol. 78, no. 1, pp. 1–3, Jan. 1950.
  • F. Pedregosa et al., “Scikit-learn: Machine Learning in Python,” Journal of Machine Learning Research, vol. 12, pp. 2825–2830, 2011.
  • L. Breiman, “Random Forests,” Machine Learning, vol. 45, no. 1, pp. 5–32, 2001.
  • T. M. Mitchell, Machine Learning. New York, NY, USA: McGraw-Hill, 1997.
  • A. Géron, Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow, 2nd ed. Sebastopol, CA, USA: O’Reilly Media, 2019.
Index Terms
Computer Science
Information Sciences
No index terms available.
Keywords

Calibration Selective Prediction Expected Calibration Error COMPAS Recidivism Trustworthy AI Temperature Scaling Abstention.

Powered by PhDFocusTM