Correlation Imputation for Single-Cell RNA-seq

Luqin Gan; Giuseppe Vinci; Genevera I Allen

doi:10.1089/cmb.2021.0403

Correlation Imputation for Single-Cell RNA-seq

J Comput Biol. 2022 May;29(5):465-482. doi: 10.1089/cmb.2021.0403. Epub 2022 Mar 21.

Authors

Luqin Gan¹, Giuseppe Vinci², Genevera I Allen^{1

3

4

5}

Affiliations

¹ Department of Statistics, Rice University, Houston, Texas, USA.
² Department of Applied and Computational Mathematics and Statistics University of Notre Dame, Notre Dame, Indiana, USA.
³ Department of Electrical and Computer Engineering and Rice University, Houston, Texas, USA.
⁴ Department of Computer Science, Rice University, Houston, Texas, USA.
⁵ Neurological Research Institute, Baylor College of Medicine, Houston, Texas, USA.

Abstract

Recent advances in single-cell RNA sequencing (scRNA-seq) technologies have yielded a powerful tool to measure gene expression of individual cells. One major challenge of the scRNA-seq data is that it usually contains a large amount of zero expression values, which often impairs the effectiveness of downstream analyses. Numerous data imputation methods have been proposed to deal with these "dropout" events, but this is a difficult task for such high-dimensional and sparse data. Furthermore, there have been debates on the nature of the sparsity, about whether the zeros are due to technological limitations or represent actual biology. To address these challenges, we propose Single-cell RNA-seq Correlation completion by ENsemble learning and Auxiliary information (SCENA), a novel approach that imputes the correlation matrix of the data of interest instead of the data itself. SCENA obtains a gene-by-gene correlation estimate by ensembling various individual estimates, some of which are based on known auxiliary information about gene expression networks. Our approach is a reliable method that makes no assumptions on the nature of sparsity in scRNA-seq data or the data distribution. By extensive simulation studies and real data applications, we demonstrate that SCENA is not only superior in gene correlation estimation, but also improves the accuracy and reliability of downstream analyses, including cell clustering, dimension reduction, and graphical model estimation to learn the gene expression network.

Keywords: auxiliary information; clustering; correlation completion; dimension reduction; ensemble learning; graphical modeling; imputation; single-cell RNA-sequencing.

Publication types

Research Support, N.I.H., Extramural
Research Support, Non-U.S. Gov't
Research Support, U.S. Gov't, Non-P.H.S.

MeSH terms

Cluster Analysis
Computer Simulation
Gene Expression Profiling*
RNA-Seq
Reproducibility of Results
Sequence Analysis, RNA / methods
Single-Cell Analysis* / methods

Grants and funding

R01 GM140468/GM/NIGMS NIH HHS/United States