Inference of differentially expressed genes using generalized linear mixed models in a pairwise fashion

PeerJ. 2023 Apr 3:11:e15145. doi: 10.7717/peerj.15145. eCollection 2023.

Abstract

Background: Technological advances involving RNA-Seq and Bioinformatics allow quantifying the transcriptional levels of genes in cells, tissues, and cell lines, permitting the identification of Differentially Expressed Genes (DEGs). DESeq2 and edgeR are well-established computational tools used for this purpose and they are based upon generalized linear models (GLMs) that consider only fixed effects in modeling. However, the inclusion of random effects reduces the risk of missing potential DEGs that may be essential in the context of the biological phenomenon under investigation. The generalized linear mixed models (GLMM) can be used to include both effects.

Methods: We present DEGRE (Differentially Expressed Genes with Random Effects), a user-friendly tool capable of inferring DEGs where fixed and random effects on individuals are considered in the experimental design of RNA-Seq research. DEGRE preprocesses the raw matrices before fitting GLMMs on the genes and the derived regression coefficients are analyzed using the Wald statistical test. DEGRE offers the Benjamini-Hochberg or Bonferroni techniques for P-value adjustment.

Results: The datasets used for DEGRE assessment were simulated with known identification of DEGs. These have fixed effects, and the random effects were estimated and inserted to measure the impact of experimental designs with high biological variability. For DEGs' inference, preprocessing effectively prepares the data and retains overdispersed genes. The biological coefficient of variation is inferred from the counting matrices to assess variability before and after the preprocessing. The DEGRE is computationally validated through its performance by the simulation of counting matrices, which have biological variability related to fixed and random effects. DEGRE also provides improved assessment measures for detecting DEGs in cases with higher biological variability. We show that the preprocessing established here effectively removes technical variation from those matrices. This tool also detects new potential candidate DEGs in the transcriptome data of patients with bipolar disorder, presenting a promising tool to detect more relevant genes.

Conclusions: DEGRE provides data preprocessing and applies GLMMs for DEGs' inference. The preprocessing allows efficient remotion of genes that could impact the inference. Also, the computational and biological validation of DEGRE has shown to be promising in identifying possible DEGs in experiments derived from complex experimental designs. This tool may help handle random effects on individuals in the inference of DEGs and presents a potential for discovering new interesting DEGs for further biological investigation.

Keywords: DEGRE package; Differentially expressed genes; Gene dispersion; Generalized linear mixed model; Preprocessing; Random effects.

Publication types

  • Research Support, Non-U.S. Gov't

MeSH terms

  • Computational Biology / methods
  • Gene Expression Profiling* / methods
  • Humans
  • Linear Models
  • Transcriptome* / genetics

Grants and funding

This work was developed in the frameworks of thematic project FAPERJ E-26/210.012/2020 and E-26/210.681/2021. Ana Tereza Ribeiro de Vasconcelos is supported by CNPq (307145/2021-2) and FAPERJ (E-26/201.046/2022). Douglas Terra Machado was supported by CAPES (88882.332653/2019-01). Yasmmin Côrtes Martins was supported by FAPERJ (E-26/202.168/2020). The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.