EpiGePT: a pretrained transformer-based language model for context-specific human epigenomics

Genome Biol. 2024 Dec 18;25(1):310. doi: 10.1186/s13059-024-03449-7.

Abstract

The inherent similarities between natural language and biological sequences have inspired the use of large language models in genomics, but current models struggle to incorporate chromatin interactions or predict in unseen cellular contexts. To address this, we propose EpiGePT, a transformer-based model designed for predicting context-specific human epigenomic signals. By incorporating transcription factor activities and 3D genome interactions, EpiGePT outperforms existing methods in epigenomic signal prediction tasks, especially in cell-type-specific long-range interaction predictions and genetic variant impacts, advancing our understanding of gene regulation. A free online prediction service is available at http://health.tsinghua.edu.cn/epigept .

Keywords: 3D genome; Epigenomics; Gene Regulation; Language model; Transformer.

MeSH terms

  • Chromatin / genetics
  • Chromatin / metabolism
  • Epigenesis, Genetic
  • Epigenomics* / methods
  • Genome, Human
  • Humans
  • Software
  • Transcription Factors / genetics
  • Transcription Factors / metabolism

Substances

  • Chromatin
  • Transcription Factors