Reproducible detection of disease-associated markers from gene expression data

BMC Med Genomics. 2016 Aug 18;9(1):53. doi: 10.1186/s12920-016-0214-5.

Abstract

Background: Detection of disease-associated markers plays a crucial role in gene screening for biological studies. Two-sample test statistics, such as the t-statistic, are widely used to rank genes based on gene expression data. However, the resultant gene ranking is often not reproducible among different data sets. Such irreproducibility may be caused by disease heterogeneity.

Results: When we divided data into two subsets, we found that the signs of the two t-statistics were often reversed. Focusing on such instability, we proposed a sign-sum statistic that counts the signs of the t-statistics for all possible subsets. The proposed method excludes genes affected by heterogeneity, thereby improving the reproducibility of gene ranking. We compared the sign-sum statistic with the t-statistic by a theoretical evaluation of the upper confidence limit. Through simulations and applications to real data sets, we show that the sign-sum statistic exhibits superior performance.

Conclusion: We derive the sign-sum statistic for getting a robust gene ranking. The sign-sum statistic gives more reproducible ranking than the t-statistic. Using simulated data sets we show that the sign-sum statistic excludes hetero-type genes well. Also for the real data sets, the sign-sum statistic performs well in a viewpoint of ranking reproducibility.

Keywords: Gene expression analysis; Genes screening; Heterogeneity; Subsampling method; Two-sample test; U-statistic.

Publication types

  • Research Support, Non-U.S. Gov't

MeSH terms

  • Biomarkers / metabolism
  • Computational Biology / methods*
  • Disease / genetics*
  • Gene Expression Profiling*
  • Humans
  • Reproducibility of Results

Substances

  • Biomarkers