Résumé
We extended our phylo-k-mer classifier to simultaneously take into account the probability and the collinearity of phylo-k-mer matches as a means to filtersequences that are actually related to a particular phylo-k-mer index. We implemented a simple scoring scheme based on entropy and accumulated probabilities. In practice, we rapidly extended this idea to allow classification into a collection of phylogenies. First, we build a phylo-k-mer index for each phylogeny. Then, we build a meta-index of phylo-k-mers selected among these indices. A similar classifier is then used to evaluate which index should be selected, then classification into the corresponding tree branches can be computed (Figure 1). This approach appears to be sufficient even when phylogenies are based on homologous sequences. For instance, it can successfully distinguish gene trees showing similar domains.