Counting pseudoalignments to novel splicing events

Luka Borozan; Francisca Rojas Ringeling; Shao-Yen Kao; Elena Nikonova; Pablo Monteagudo-Mesas; Domagoj Matijević; Maria L Spletter; Stefan Canzar

doi:10.1093/bioinformatics/btad419

Counting pseudoalignments to novel splicing events

Bioinformatics. 2023 Jul 1;39(7):btad419. doi: 10.1093/bioinformatics/btad419.

Authors

Luka Borozan¹, Francisca Rojas Ringeling^{2

3}, Shao-Yen Kao⁴, Elena Nikonova⁴, Pablo Monteagudo-Mesas², Domagoj Matijević¹, Maria L Spletter^{4

5}, Stefan Canzar^{2

3

6}

Affiliations

¹ Department of Mathematics, Josip Juraj Strossmayer University of Osijek, Osijek 31000, Croatia.
² Gene Center, Ludwig-Maximilians-Universität München, Munich 81377, Germany.
³ Huck Institutes of the Life Sciences, The Pennsylvania State University, University Park, PA 16802, United States.
⁴ Biomedical Center, Department of Physiological Chemistry, Ludwig-Maximilians-Universität München, Planegg-Martinsried 82152, Germany.
⁵ School of Science and Engineering, Division of Biological & Biomedical Systems, University of Missouri Kansas City, Kansas City, MO 64110, United States.
⁶ Department of Computer Science and Engineering, The Pennsylvania State University, University Park, PA 16802, United States.

Abstract

Motivation: Alternative splicing (AS) of introns from pre-mRNA produces diverse sets of transcripts across cell types and tissues, but is also dysregulated in many diseases. Alignment-free computational methods have greatly accelerated the quantification of mRNA transcripts from short RNA-seq reads, but they inherently rely on a catalog of known transcripts and might miss novel, disease-specific splicing events. By contrast, alignment of reads to the genome can effectively identify novel exonic segments and introns. Event-based methods then count how many reads align to predefined features. However, an alignment is more expensive to compute and constitutes a bottleneck in many AS analysis methods.

Results: Here, we propose fortuna, a method that guesses novel combinations of annotated splice sites to create transcript fragments. It then pseudoaligns reads to fragments using kallisto and efficiently derives counts of the most elementary splicing units from kallisto's equivalence classes. These counts can be directly used for AS analysis or summarized to larger units as used by other widely applied methods. In experiments on synthetic and real data, fortuna was around 7× faster than traditional align and count approaches, and was able to analyze almost 300 million reads in just 15 min when using four threads. It mapped reads containing mismatches more accurately across novel junctions and found more reads supporting aberrant splicing events in patients with autism spectrum disorder than existing methods. We further used fortuna to identify novel, tissue-specific splicing events in Drosophila.

Availability and implementation: fortuna source code is available at https://github.com/canzarlab/fortuna.

Publication types

Research Support, Non-U.S. Gov't

MeSH terms

Alternative Splicing
Autism Spectrum Disorder*
Humans
RNA Splicing
Sequence Analysis, RNA / methods
Software