Adjacency-constrained hierarchical clustering of a band similarity matrix with application to genomics

Christophe Ambroise; Alia Dehman; Pierre Neuvial; Guillem Rigaill; Nathalie Vialaneix

doi:10.1186/s13015-019-0157-4

Adjacency-constrained hierarchical clustering of a band similarity matrix with application to genomics

Algorithms Mol Biol. 2019 Nov 15:14:22. doi: 10.1186/s13015-019-0157-4. eCollection 2019.

Authors

Christophe Ambroise¹, Alia Dehman², Pierre Neuvial³, Guillem Rigaill^{1

4}, Nathalie Vialaneix⁵

Affiliations

¹ 1Laboratoire de Mathématiques et Modélisation d'Evry, UMR CNRS 8071, Université d'Evry Val d'Essonne, 23 boulevard de France, 91037 Evry, France.
² Hyphen-stat, 195 Route d'Espagne, 31036 Toulouse, France.
³ 3Institut de Mathématiques de Toulouse, UMR5219 CNRS, Université de Toulouse, UPS IMT, 31062 Toulouse Cedex 9, France.
⁴ 4Institute of Plant Sciences Paris Saclay IPS2, CNRS, INRA, Gif sur Yvette, France.
⁵ MIAT, Université de Toulouse, INRA, Castanet-Tolosan, France.

Abstract

Background: Genomic data analyses such as Genome-Wide Association Studies (GWAS) or Hi-C studies are often faced with the problem of partitioning chromosomes into successive regions based on a similarity matrix of high-resolution, locus-level measurements. An intuitive way of doing this is to perform a modified Hierarchical Agglomerative Clustering (HAC), where only adjacent clusters (according to the ordering of positions within a chromosome) are allowed to be merged. But a major practical drawback of this method is its quadratic time and space complexity in the number of loci, which is typically of the order of $10^{4}$ to $10^{5}$ for each chromosome.

Results: By assuming that the similarity between physically distant objects is negligible, we are able to propose an implementation of adjacency-constrained HAC with quasi-linear complexity. This is achieved by pre-calculating specific sums of similarities, and storing candidate fusions in a min-heap. Our illustrations on GWAS and Hi-C datasets demonstrate the relevance of this assumption, and show that this method highlights biologically meaningful signals. Thanks to its small time and memory footprint, the method can be run on a standard laptop in minutes or even seconds.

Availability and implementation: Software and sample data are available as an R package, adjclust, that can be downloaded from the Comprehensive R Archive Network (CRAN).

Keywords: Adjacency constraint; Genome-Wide Association Studies and Hi-C; Hierarchical agglomerative clustering; Min heap; Segmentation; Similarity; Ward’s linkage.