Seq2science: an end-to-end workflow for functional genomics analysis

PeerJ. 2023 Nov 15:11:e16380. doi: 10.7717/peerj.16380. eCollection 2023.

Abstract

Sequencing databases contain enormous amounts of functional genomics data, making them an extensive resource for genome-scale analysis. Reanalyzing publicly available data, and integrating it with new, project-specific data sets, can be invaluable. With current technologies, genomic experiments have become feasible for virtually any species of interest. However, using and integrating this data comes with its challenges, such as standardized and reproducible analysis. Seq2science is a multi-purpose workflow that covers preprocessing, quality control, visualization, and analysis of functional genomics sequencing data. It facilitates the downloading of sequencing data from all major databases, including NCBI SRA, EBI ENA, DDBJ, GSA, and ENCODE. Furthermore, it automates the retrieval of any genome assembly available from Ensembl, NCBI, and UCSC. It has been tested on a variety of species, and includes diverse workflows such as ATAC-, RNA-, and ChIP-seq. It consists of both generic as well as advanced steps, such as differential gene expression or peak accessibility analysis and differential motif analysis. Seq2science is built on the Snakemake workflow language and thus can be run on a range of computing infrastructures. It is available at https://github.com/vanheeringen-lab/seq2science.

Keywords: ATAC-seq; ChIP-seq; NGS; RNA-seq; Workflow.

MeSH terms

  • Chromatin Immunoprecipitation Sequencing
  • Genomics
  • High-Throughput Nucleotide Sequencing*
  • Software*
  • Workflow

Grants and funding

This work was supported by the Netherlands Organization for Scientific Research (NWO Grant 016.Vidi.189.081). The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.