Development of a Structured Query Language and Natural Language Processing Algorithm to Identify Lung Nodules in a Cancer Centre

Benjamin Hunter; Sara Reis; Des Campbell; Sheila Matharu; Prashanthi Ratnakumar; Luca Mercuri; Sumeet Hindocha; Hardeep Kalsi; Erik Mayer; Ben Glampson; Emily J Robinson; Bisan Al-Lazikani; Lisa Scerri; Susannah Bloch; Richard Lee

doi:10.3389/fmed.2021.748168

Development of a Structured Query Language and Natural Language Processing Algorithm to Identify Lung Nodules in a Cancer Centre

Front Med (Lausanne). 2021 Nov 4:8:748168. doi: 10.3389/fmed.2021.748168. eCollection 2021.

Authors

Benjamin Hunter^{1

2}, Sara Reis¹, Des Campbell¹, Sheila Matharu¹, Prashanthi Ratnakumar³, Luca Mercuri⁴, Sumeet Hindocha^{1

2}, Hardeep Kalsi^{1

2}, Erik Mayer^{2

4}, Ben Glampson⁴, Emily J Robinson⁵, Bisan Al-Lazikani⁶, Lisa Scerri¹, Susannah Bloch³, Richard Lee^{1

7

8}

Affiliations

¹ The Royal Marsden National Health Service (NHS) Foundation Trust, Lung Unit, London, United Kingdom.
² Department of Surgery and Cancer, Imperial College London, London, United Kingdom.
³ Imperial College Healthcare Trust, Respiratory Medicine, London, United Kingdom.
⁴ Imperial College Healthcare National Health Service (NHS) Trust, Imperial Clinical Analytics, Research and Evaluation, London, United Kingdom.
⁵ The Royal Marsden National Health Service (NHS) Foundation Trust, Royal Marsden Clinical Trials Unit, London, United Kingdom.
⁶ The Institute for Cancer Research, Computational Biology and Chromogenetics, London, United Kingdom.
⁷ Imperial College London, National Heart and Lung Institute, London, United Kingdom.
⁸ The Institute for Cancer Research, Early Diagnosis and Detection, Genetics and Epidemiology, London, United Kingdom.

Abstract

Importance: The stratification of indeterminate lung nodules is a growing problem, but the burden of lung nodules on healthcare services is not well-described. Manual service evaluation and research cohort curation can be time-consuming and potentially improved by automation. Objective: To automate lung nodule identification in a tertiary cancer centre. Methods: This retrospective cohort study used Electronic Healthcare Records to identify CT reports generated between 31st October 2011 and 24th July 2020. A structured query language/natural language processing tool was developed to classify reports according to lung nodule status. Performance was externally validated. Sentences were used to train machine-learning classifiers to predict concerning nodule features in 2,000 patients. Results: 14,586 patients with lung nodules were identified. The cancer types most commonly associated with lung nodules were lung (39%), neuro-endocrine (38%), skin (35%), colorectal (33%) and sarcoma (33%). Lung nodule patients had a greater proportion of metastatic diagnoses (45 vs. 23%, p < 0.001), a higher mean post-baseline scan number (6.56 vs. 1.93, p < 0.001), and a shorter mean scan interval (4.1 vs. 5.9 months, p < 0.001) than those without nodules. Inter-observer agreement for sentence classification was 0.94 internally and 0.98 externally. Sensitivity and specificity for nodule identification were 93 and 99% internally, and 100 and 100% at external validation, respectively. A linear-support vector machine model predicted concerning sentence features with 94% accuracy. Conclusion: We have developed and validated an accurate tool for automated lung nodule identification that is valuable for service evaluation and research data acquisition.

Keywords: informatics; lung nodule; machine learning; natural language processing (NLP); structured query language (SQL).