Enabling integrative genomic analysis of high-impact human diseases through text mining

Joel Dudley; Atul J Butte

Enabling integrative genomic analysis of high-impact human diseases through text mining

Pac Symp Biocomput. 2008:580-91.

Authors

Joel Dudley¹, Atul J Butte

Affiliation

¹ Stanford Medical Informatics, Departments of Medicine and Pediatrics, Stanford University School of Medicine, Stanford, CA 94305-5479, USA.

PMID: 18229717
PMCID: PMC2735266

Abstract

Our limited ability to perform large-scale translational discovery and analysis of disease characterizations from public genomic data repositories remains a major bottleneck in efforts to translate genomics experiments to medicine. Through comprehensive, integrative genomic analysis of all available human disease characterizations we gain crucial insight into the molecular phenomena underlying pathogenesis as well as intra- and inter-disease differentiation. Such knowledge is crucial in the development of improved clinical diagnostics and the identification of molecular targets for novel therapeutics. In this study we build on our previous work to realize the next important step in large-scale translational discovery and analysis, which is to automatically identify those genomic experiments in which a disease state is compared to a normal control state. We present an automated text mining method that employs Natural Language Processing (NLP) techniques to automatically identify disease-related experiments in the NCBI Gene Expression Omnibus (GEO) that include measurements for both disease and normal control states. In this manner, we find that 62% of disease-related experiments contain sample subsets that can be automatically identified as normal controls. Furthermore, we calculate that the identified experiments characterize diseases that contribute to 30% of all human disease-related mortality in the United States. This work demonstrates that we now have the necessary tools and methods to initiate large-scale translational bioinformatics inquiry across the broad spectrum of high-impact human disease.

Publication types

Evaluation Study
Research Support, N.I.H., Extramural
Research Support, Non-U.S. Gov't

MeSH terms

Computational Biology*
Databases, Genetic
Genetic Diseases, Inborn / genetics*
Genomics* / statistics & numerical data
Humans
Information Storage and Retrieval*
Oligonucleotide Array Sequence Analysis / statistics & numerical data
PubMed
Unified Medical Language System

Abstract

Publication types

MeSH terms

Grants and funding