Predicting human health from biofluid-based metabolomics using machine learning

Ethan D Evans; Claire Duvallet; Nathaniel D Chu; Michael K Oberst; Michael A Murphy; Isaac Rockafellow; David Sontag; Eric J Alm

doi:10.1038/s41598-020-74823-1

Predicting human health from biofluid-based metabolomics using machine learning

Sci Rep. 2020 Oct 19;10(1):17635. doi: 10.1038/s41598-020-74823-1.

Authors

Ethan D Evans¹, Claire Duvallet^{1

2}, Nathaniel D Chu¹, Michael K Oberst³, Michael A Murphy^{1

3}, Isaac Rockafellow^{1

4}, David Sontag⁵, Eric J Alm⁶

Affiliations

¹ Department of Biological Engineering, Massachusetts Institute of Technology, Cambridge, MA, 02139, USA.
² Biobot Analytics, Somerville, MA, 02143, USA.
³ CSAIL, Massachusetts Institute of Technology, Cambridge, MA, 02139, USA.
⁴ Superpedestrian, Cambridge, MA, 02139, USA.
⁵ CSAIL, Massachusetts Institute of Technology, Cambridge, MA, 02139, USA. [email protected].
⁶ Department of Biological Engineering, Massachusetts Institute of Technology, Cambridge, MA, 02139, USA. [email protected].

Abstract

Biofluid-based metabolomics has the potential to provide highly accurate, minimally invasive diagnostics. Metabolomics studies using mass spectrometry typically reduce the high-dimensional data to only a small number of statistically significant features, that are often chemically identified-where each feature corresponds to a mass-to-charge ratio, retention time, and intensity. This practice may remove a substantial amount of predictive signal. To test the utility of the complete feature set, we train machine learning models for health state-prediction in 35 human metabolomics studies, representing 148 individual data sets. Models trained with all features outperform those using only significant features and frequently provide high predictive performance across nine health state categories, despite disparate experimental and disease contexts. Using only non-significant features it is still often possible to train models and achieve high predictive performance, suggesting useful predictive signal. This work highlights the potential for health state diagnostics using all metabolomics features with data-driven analysis.

Publication types

Research Support, Non-U.S. Gov't

MeSH terms

Databases, Factual
Health Status
Humans
Machine Learning*
Metabolomics / methods*
Models, Theoretical*