Résumé
Understanding genotype-phenotype relationships is one of the most important research axes in agronomy. New challenges involve elucidating the relationships among various molecular components responsible for phenome expression. It appears that these challenges can only be addressed effectively by integrating information across multiple biological scales into a global model through a systemic approach, thus facilitating a comprehensive understanding of the true functioning of biological systems. Recent high-throughput analysis technologies can only partially capture this complexity. While these technologies enable continuous advancements in data acquisition, reflection on centralizing data into dedicated platforms and standardizing formats should allow more effective data integration, thereby improving overall knowledge. Indeed, current knowledge remains fragmented, hindering the elucidation of molecular mechanisms governing complex phenotypic traits.My research project aligns with these reflections, addressing the following issue: how to structure and manage the complexity of biological data to extract meaningful knowledge capable of identifying molecular mechanisms controlling plant phenotype expression.Our hypothesis is that creating data-driven knowledge graphs will facilitate the formulation of research hypotheses linking genotype to phenotype. Using rice as a model, the aim will be to build molecular interaction networks from scattered data to identify key genes for plant improvement. Various approaches will be leveraged regarding data integration, knowledge enrichment, and the design of knowledge graphs.Within this framework, a first approach will involve dynamically transforming and integrating relevant data into a knowledge base to enhance their analytical usability algorithmically. A second approach will propose new methods for knowledge enrichment. Initially, semantic annotation methods will be utilized, followed by the development of new data-linking methods to uncover novel relationships between generated graphs, thus yielding an interaction network conducive to discovering new insights. Finally, to enable efficient information retrieval, several candidate-gene prioritization algorithms and methods will be evaluated and proposed within available knowledge graphs.