Abstract
Over the last few decades, thanks to dramatic advances in high-throughput sequencing technologies, more and more genomes were assembled. By comparing multiple genomes within the same species, it became increasingly clear that a single reference genome cannot capture all the diversity within a species. Huge variations in the number of genes, and more broadly in structural variations, between individuals were observed. This led to a progressive paradigm shift from the genome towards the pangenome. This latter refers to the set of all sequences (genic or not) present in a given population, including (i) the core genome containing sequences present in all individuals and (ii) the dispensable genome composed of sequences shared by a subset of individuals.The main objectives of this PhD project were to fully explore the genetic diversity within a crop, the African rice, and how the diversity of this species was reshaped during its domestication. The two main questions we wanted to specifically address were which roles do these structural variations play in gene composition and adaptation, and how evolutionary forces have shaped (pan)genome organization and dynamics.First, we performed a review of the current state-of-the-art on the emerging concept of pangenome, to identify challenges, opportunities and limitations. We presented how this new methodological approach, combined with the high-throughput sequencing technologies, offers unprecedented opportunities to investigate genomic variability, thus providing an alternative way to study genome organization and dynamics.Next, we developed frangiPANe, a tool for creating a eukaryotic pangenome from multiple, individual short-read-based genome sequencing. Our method was validated using the whole genome resequencing data from 248 African rice accessions, both cultivated and wild, as a proof-of-concept to build the first large panreference for African rice. We identified an average of 8Mb of new sequences per individual, absent from the reference genome, i.e. a total of 513.5 Mb new sequence across all individuals. For a reference genome size of 350Mb, we added 60% more sequences, represented in 484,394 non redundant contigs. Despite the high content of transposable elements, we were able to anchor at an unique position on the reference genome, 31.5% of the contigs.Finally, we then explored pangenome diversity during African rice domestication. We annotated 22,765 genes across all individuals, increasing the number of genes from 40,553 in the reference by 56%. We showed the relative presence of the 484,394 contigs recapitulate the relationship between individuals previously determined using single nucleotide polymorphism, and consequently also reflects the evolutionary history of the cultivated and wild species. We thus used an innovative approach to assess selection in these 483,394 structural variations, and identify 683 candidate genes as under selection. Interestingly, we were able to find selection in the PROG1 gene, a gene selected during African and Asian rice domestication and associated with a major deletion during African rice domestication. We also showed domestication was associated with a loss of genes after the initial bottleneck, a likely consequence of an increase of drift during the domestication process.Altogether, our approaches help to reshape our thinking about the consequences of domestication on plant diversity, associated both to a loss of diversity and genes.