Comparison of Machine-Learning and Deep-Learning Methods for the Prediction of Osteoradionecrosis Resulting From Head and Neck Cancer Radiation Therapy

Brandon Reber; Lisanne Van Dijk; Brian Anderson; Abdallah Sherif Radwan Mohamed; Clifton Fuller; Stephen Lai; Kristy Brock

doi:10.1016/j.adro.2022.101163

Comparison of Machine-Learning and Deep-Learning Methods for the Prediction of Osteoradionecrosis Resulting From Head and Neck Cancer Radiation Therapy

Adv Radiat Oncol. 2022 Dec 27;8(4):101163. doi: 10.1016/j.adro.2022.101163. eCollection 2023 Jul-Aug.

Authors

Brandon Reber¹, Lisanne Van Dijk^{1

2}, Brian Anderson^{1

3}, Abdallah Sherif Radwan Mohamed¹, Clifton Fuller¹, Stephen Lai¹, Kristy Brock¹

Affiliations

¹ Department of Imaging Physics, The University of Texas MD Anderson Cancer Center, Houston, Texas.
² University of Groningen, Groningen, Netherlands.
³ University of California, San Diego, San Diego, California.

Abstract

Purpose: Deep-learning (DL) techniques have been successful in disease-prediction tasks and could improve the prediction of mandible osteoradionecrosis (ORN) resulting from head and neck cancer (HNC) radiation therapy. In this study, we retrospectively compared the performance of DL algorithms and traditional machine-learning (ML) techniques to predict mandible ORN binary outcome in an extensive cohort of patients with HNC.

Methods and materials: Patients who received HNC radiation therapy at the University of Texas MD Anderson Cancer Center from 2005 to 2015 were identified for the ML (n = 1259) and DL (n = 1236) studies. The subjects were followed for ORN development for at least 12 months, with 173 developing ORN and 1086 having no evidence of ORN. The ML models used dose-volume histogram parameters to predict ORN development. These models included logistic regression, random forest, support vector machine, and a random classifier reference. The DL models were based on ResNet, DenseNet, and autoencoder-based architectures. The DL models used each participant's dose cropped to the mandible. The effect of increasing the amount of available training data on the DL models' prediction performance was evaluated by training the DL models using increasing ratios of the original training data.

Results: The F1 score for the logistic regression model, the best-performing ML model, was 0.3. The best-performing ResNet, DenseNet, and autoencoder-based models had F1 scores of 0.07, 0.14, and 0.23, respectively, whereas the random classifier's F1 score was 0.17. No performance increase was apparent when we increased the amount of training data available for DL model training.

Conclusions: The ML models had superior performance to their DL counterparts. The lack of improvement in DL performance with increased training data suggests that either more data are needed for appropriate DL model construction or that the image features used in DL models are not suitable for this task.