Résumé
Comparing biological samples through sequencing is a core task in bioinformatics analyses such as variant detection, differential expression analysis, and epigenetic peak calling. Standard approaches typically rely on mapping newly sequenced reads to a reference genome. To avoid mapping and reference biases, k-mer-based approaches have been proposed as an alternative. Using this paradigm, our tool kdiff identifies genomic regions containing k-mers with differential abundances between samples. We demonstrate that our method effectively detects copy number variants in cancer genomes and remains robust against reference genome misassemblies. Additionally, we illustrate its utility in confirming telomere locations in noisy nanopore sequencing data. Our work demonstrates that alignment-free approaches can provide results comparable to standard alignment-based methods, while reducing the reference bias and significantly improving computational efficiency by leveraging fast k-mer counting tools.
[Display omitted]
•Alignment-free method kdiff finds differences between unassembled sequenced samples•It is robust to low-quality references and low-quality/low-coverage sequencing data•Kdiff is significantly faster than alignment-based methods
Techniques in genetics; Biocomputational method; Genomic analysis; Sequence analysis