Résumé
The huge body of publicly available RNA-seq libraries is a treasure of functional informations allowing to explore RNAs tissue expression by quantification of known RNA species or emerging novel transcript variants. However, transcript quantification using classical approaches relies on alignment methods that require a lot of computational resources, long processing time and produce possible bias due to alignment issues. Such whole transcriptome approaches can not be easily adapted to the quantification of small sets of candidate transcripts in large datasets. Recent studies have demonstrated that k-mer decomposition constitutes a new way to process RNA-Seq data for the identification of transcriptional signatures as k-mers (or tags) and can be used to quantify gene expression in a more specific and less resource-consuming way than classicalapproaches. However, applying this method to a candidate gene approach will rely on the high specificity of the k-mers set that will be quantified. Here, we present KmerExploR that includes: i/ a specific k-mers design, based on the decomposition of transcript sequences into k-mers, ii/ a subset selection of these k-mers, regarding their specificity into the reference genome and transcriptome, iii/ a counting step of the selected k-mers into RNA-seq datasets.We propose to use our strategy to set-up a pipeline for RNA-seq data quality analysis. Indeed, using well defined sets of k-mers, we are able to predict metadata from public RNA-Seq data such as library orientation, sample gender, Mycoplasma contamination or RNA ribodepletion usage. Finally, we show that k-mer analysis can also be used to test known genomic and transcriptomic modifications (mutations, splice events, fusiongenes, etc...) as well as for the discovery of new ones.