Abstract
Proteins are macromolecules at the base of the constitution of the living world. Certain regions within proteins have the ability to aggregate and form fibrils. These regions are called "amyloids" and are correlated with the onset of age-related diseases and neuropathologies. We can cite among the best known: Alzheimer’s disease, Huntington's disease, Parkinson's disease or type II diabetes. Knowing that the life expectancy increases, it is certain that the prevalence of these age-related diseases will rise in the years to come. Recent advances in technology have led to the publication of many protein sequences, structures and functions. This explosion of data has led to the emergence of a numerous bioinformatics resources and tools, notably for the prediction and analysis of amyloidogenic regions.It is in this context that we have developed TAPASS (Tool for Annotation of Protein Amyloidogenicity in the context of other Structural States), a tool for predicting amyloidogenic regions in protein sequences, as well as various other structural states. This pipeline includes ten prediction programs of different structural states of proteins : amyloidogenic regions, structured regions, unstructured regions, transmembrane regions, signal peptides and SLiMs (Short Linear Motifs). It allows us to efficiently detect potential amyloidogenic regions in the structural context of the protein. TAPASS is freely accessible all users via a web interface. Thanks to the execution speed, its use on large data sets for meta-analysis is also possible.In a second time, this thesis focused on the fundamental study of proteins amyloidogenicity. To do so, we build a database from the results of TAPASS execution on 76 reference proteomes representing the diversity of the three kingdoms of life (archaea, bacteria and eukaryote). The database allowed us to study the amyloidogenic regions (AR) within the unstructured regions, also called exposed amyloidogenic regions (EAR), presenting a higher potential to aggregate. This allowed us to establish several novel correlations on the amyloidogenicity of different species. We found that prokaryotes have a higher proportion of AR containing proteins, while eukaryotes have more EAR containing proteins. Thermophilic prokaryotes, species living in high temperature environments, have fewer AR and EAR compared to mesophilic prokaryotes. We have also established that EAR containing proteins are preferentially localized in the cell nucleus of eukaryotes, but tend to be more secreted in prokaryotes. In addition, the more the length of proteins increases, the less they contain ARs, but this decrease is not observed with EARs. We observed a significant enrichment of EAR containing proteins with SLiMs in comparison to intrinsically disordered proteins without EAR. Finally, proteins with high abundance have less EARs when compare to proteins with low abundance. Our analysis also showed that essential proteins have less EARs in comparison to non-essential ones.This thesis work also contributed to the development of AmyloComp. This program able to predict the co-aggregation between two different amyloidogenic regions. It is based on the detection of B-arches with the help of the improved version of ArchCandy 2.0.