Ten common issues with reference sequence databases and how to mitigate them

Front Bioinform. 2024 Mar 15:4:1278228. doi: 10.3389/fbinf.2024.1278228. eCollection 2024.

Abstract

Metagenomic sequencing has revolutionized our understanding of microbiology. While metagenomic tools and approaches have been extensively evaluated and benchmarked, far less attention has been given to the reference sequence database used in metagenomic classification. Issues with reference sequence databases are pervasive. Database contamination is the most recognized issue in the literature; however, it remains relatively unmitigated in most analyses. Other common issues with reference sequence databases include taxonomic errors, inappropriate inclusion and exclusion criteria, and sequence content errors. This review covers ten common issues with reference sequence databases and the potential downstream consequences of these issues. Mitigation measures are discussed for each issue, including bioinformatic tools and database curation strategies. Together, these strategies present a path towards more accurate, reproducible and translatable metagenomic sequencing.

Keywords: database; metagenomic; reference; sequence; taxonomy.

Publication types

  • Review

Grants and funding

The author(s) declare that no financial support was received for the research, authorship, and/or publication of this article.