Classifying Cyber-Risky Clinical Notes by Employing Natural Language Processing

Schmeelk, Suzanna; Dogo, Martins Samuel; Peng, Yifan; Patra, Braja Gopal

doi:10.24251/HICSS.2022.505

Computer Science > Computation and Language

arXiv:2203.12781 (cs)

[Submitted on 24 Mar 2022]

Title:Classifying Cyber-Risky Clinical Notes by Employing Natural Language Processing

Authors:Suzanna Schmeelk, Martins Samuel Dogo, Yifan Peng, Braja Gopal Patra

View PDF

Abstract:Clinical notes, which can be embedded into electronic medical records, document patient care delivery and summarize interactions between healthcare providers and patients. These clinical notes directly inform patient care and can also indirectly inform research and quality/safety metrics, among other indirect metrics. Recently, some states within the United States of America require patients to have open access to their clinical notes to improve the exchange of patient information for patient care. Thus, developing methods to assess the cyber risks of clinical notes before sharing and exchanging data is critical. While existing natural language processing techniques are geared to de-identify clinical notes, to the best of our knowledge, few have focused on classifying sensitive-information risk, which is a fundamental step toward developing effective, widespread protection of patient health information. To bridge this gap, this research investigates methods for identifying security/privacy risks within clinical notes. The classification either can be used upstream to identify areas within notes that likely contain sensitive information or downstream to improve the identification of clinical notes that have not been entirely de-identified. We develop several models using unigram and word2vec features with different classifiers to categorize sentence risk. Experiments on i2b2 de-identification dataset show that the SVM classifier using word2vec features obtained a maximum F1-score of 0.792. Future research involves articulation and differentiation of risk in terms of different global regulatory requirements.

Comments:	7 pages, 4 figures, published in Proceedings of the 55th Hawaii International Conference on System Sciences
Subjects:	Computation and Language (cs.CL); Computers and Society (cs.CY)
Cite as:	arXiv:2203.12781 [cs.CL]
	(or arXiv:2203.12781v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2203.12781
Related DOI:	https://doi.org/10.24251/HICSS.2022.505

Submission history

From: Martins Samuel Dogo [view email]
[v1] Thu, 24 Mar 2022 00:36:59 UTC (724 KB)

Computer Science > Computation and Language

Title:Classifying Cyber-Risky Clinical Notes by Employing Natural Language Processing

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Classifying Cyber-Risky Clinical Notes by Employing Natural Language Processing

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators