Validity and reliability of Brier scoring for assessment of probabilistic diagnostic reasoning

Nathan Stehouwer; Anastasia Rowland-Seymour; Larry Gruppen; Jeffrey M Albert; Kelli Qua

doi:10.1515/dx-2023-0109

Validity and reliability of Brier scoring for assessment of probabilistic diagnostic reasoning

Diagnosis (Berl). 2024 Oct 16. doi: 10.1515/dx-2023-0109. Online ahead of print.

Authors

Nathan Stehouwer^{1

2}, Anastasia Rowland-Seymour^{2

3}, Larry Gruppen⁴, Jeffrey M Albert², Kelli Qua²

Affiliations

¹ University Hospitals Cleveland Medical Center and Rainbow Babies & Children's Hospital, Cleveland, OH, USA.
² Case Western Reserve University School of Medicine, Cleveland, OH, USA.
³ MetroHealth Medical Center, Cleveland, OH, USA.
⁴ University of Michigan Medical School, Ann Arbor, MI, USA.

PMID: 39402892
DOI: 10.1515/dx-2023-0109

Abstract

Objectives: Educators need tools for the assessment of clinical reasoning that reflect the ambiguity of real-world practice and measure learners' ability to determine diagnostic likelihood. In this study, the authors describe the use of the Brier score to assess and provide feedback on the quality of probabilistic diagnostic reasoning.

Methods: The authors describe a novel format called Diagnostic Forecasting (DxF), in which participants read a brief clinical case and assign a probability to each item on a differential diagnosis, order tests and select a final diagnosis. DxF was piloted in a cohort of senior medical students. DxF evaluated students' answers with Brier scores, which compare probabilistic forecasts with case outcomes. The validity of Brier scores in DxF was assessed by comparison to subsequent decision-making in the game environment of DxF, as well as external criteria including medical knowledge tests and performance on clinical rotations.

Results: Brier scores were statistically significantly correlated with diagnostic accuracy (95 % CI -4.4 to -0.44) and with mean scores on the National Board of Medical Examiners (NBME) shelf exams (95 % CI -474.6 to -225.1). Brier scores did not correlate with clerkship grades or performance on a structured clinical skills exam. Reliability as measured by within-student correlation was low.

Conclusions: Brier scoring showed evidence for validity as a measurement of medical knowledge and predictor of clinical decision-making. Further work must evaluated the ability of Brier scores to predict clinical and workplace-based outcomes, and develop reliable approaches to measuring probabilistic reasoning.

Keywords: assessment; diagnostic reasoning; probabilistic reasoning; uncertainty.