Logo image
Zero-shot deep learning for the annotation of unknown eDNA sequences using co-occurrences and phylogenetic embeddings
Article de revue   Open Access   Avec comité de lecture

Zero-shot deep learning for the annotation of unknown eDNA sequences using co-occurrences and phylogenetic embeddings

Steven Stalder, Théophile Sanchez, Michele Volpi, Stéphanie Manel, David Mouillot, Arnaud Auber, Morgane Bruno, Virginie Marques, Camille Albouy et Loïc Pellissier
PLoS Computational Biology, Vol.21(12)
2025
PMID: 41417866

Résumé

The advent of environmental DNA (eDNA) metabarcoding marks a transformative era in large-scale biodiversity monitoring. However, the analysis of eDNA datasets is limited by incomplete reference databases and the increasing volume of data requiring processing from raw sequences to annotated taxonomic lists. To curate taxonomic lists from eDNA analysis, geographic constraints are used by expert in postanalysis, which may introduce potential biases in assignments. Instead of relying on expert intervention, a combination of taxonomic and geographic co-occurrences could be directly integrated into machine learning to automatize and improve taxonomic annotation. Here, we introduce a deep learning approach applied to the taxonomic assignment of eDNA sequences, which leverages a species reference database, species co-occurrence data, and a phylogeny to enhance annotation directly from raw sequences. The phylogeny provides the structure to the network's embedding space in which DNA sequences are placed utilizing an artificial neural network (ANN). We train an additional ANN from the phylogenetic embedding and co-occurrence species data to learn coherent species combinations from the whole collection of eDNA sequences, as opposed to single sequences only. When applied directly to the raw sequences, this method correctly predicts unseen species (i.e., those not contained in the reference database), out of more than 31,000 possibilities, in about 24% of the tested cases by relying on phylogenetic embeddings and geographic modulation. The trained ANNs discern species relationships accurately from the raw data, which facilitates the process of associating sequences with application, which we have developed to easily run our pipeline without prior setup or computational resources. Trawl contents are publicly available for the North East Atlantic (DATRAS database; https://datras.ices.

Fichiers et liens (2)

url
Find in HALAfficher
url
https://doi.org/10.1371/journal.pcbi.1013776Afficher
Published (Version of record) Ouvrir

Indicateurs

1 Consultations de la notice

Détails

Logo image