When conducting database studies, researchers sometimes use an algorithm known as "case definition," "outcome definition," or "computable phenotype" to identify the outcome of interest. Generally, algorithms are created by combining multiple variables and codes, and we need to select the most appropriate one to apply to the database study. Validation studies compare algorithms with the gold standard and calculate indicators such as sensitivity and specificity to assess their validities. As the indicators are calculated for each algorithm, selecting an algorithm is equivalent to choosing a pair of sensitivity and specificity. Therefore, receiver operating characteristic curves can be utilized, and two intuitive criteria are commonly used. However, neither was conceived to reduce the biases of effect measures (e.g., risk difference and risk ratio), which are important in database studies. In this study, we evaluated two existing criteria from perspectives of the biases and found that one of them, called the Youden index always minimizes the bias of the risk difference regardless of the true incidence proportions under nondifferential outcome misclassifications. However, both criteria may lead to inaccurate estimates of absolute risks, and such property is undesirable in decision-making. Therefore, we propose a new criterion based on minimizing the sum of the squared biases of absolute risks to estimate them more accurately. Subsequently, we apply all criteria to the data from the actual validation study on postsurgical infections and present the results of a sensitivity analysis to examine the robustness of the assumption our proposed criterion requires.
Copyright © 2024 The Author(s). Published by Wolters Kluwer Health, Inc.