Extracting subject demographic information from abstracts of randomized clinical trial reports

Rong Xu; Yael Garten; Kaustubh S Supekar; Amar K Das; Russ B Altman; Alan M Garber

Extracting subject demographic information from abstracts of randomized clinical trial reports

Stud Health Technol Inform. 2007;129(Pt 1):550-4.

Authors

Rong Xu¹, Yael Garten, Kaustubh S Supekar, Amar K Das, Russ B Altman, Alan M Garber

Affiliation

¹ Biomedical Informatics Training Program, Stanford University School of Medicine, USA.

PMID: 17911777

Abstract

In order to make more informed healthcare decisions, consumers need information systems that deliver accurate and reliable information about their illnesses and potential treatments. Reports of randomized clinical trials (RCTs) provide reliable medical evidence about the efficacy of treatments. Current methods to access, search for, and retrieve RCTs are keyword-based, time-consuming, and suffer from poor precision. Personalized semantic search and medical evidence summarization aim to solve this problem. The performance of these approaches may improve if they have access to study subject descriptors (e.g. age, gender, and ethnicity), trial sizes, and diseases/symptoms studied. We have developed a novel method to automatically extract such subject demographic information from RCT abstracts. We used text classification augmented with a Hidden Markov Model to identify sentences containing subject demographics, and subsequently these sentences were parsed using Natural Language Processing techniques to extract relevant information. Our results show accuracy levels of 82.5%, 92.5%, and 92.0% for extraction of subject descriptors, trial sizes, and diseases/symptoms descriptors respectively.

Publication types

Research Support, N.I.H., Extramural
Research Support, U.S. Gov't, Non-P.H.S.

MeSH terms

Abstracting and Indexing*
Information Storage and Retrieval / methods*
Markov Chains
Natural Language Processing*
Randomized Controlled Trials as Topic

Abstract

Publication types

MeSH terms

Grants and funding