Abstract
The time has come for machines to overpower human comprehension. At an era where most things become virtualized, algorithms rules our lives and the accumulated knowledge from advances in science enabled computer to accurately predict protein structures, the simulation of interactions between biochemical entities still remain a challenge. Since the beginning of human kind, the quest of enlightenment has always been of major interest. By describing and understanding its surroundings, from stars to atoms, researches gave permission to develop mathematical models, new methodologies and experimental techniques susceptible of performing precise measurements of the biological systems and grasp a picture of life.Driven by thermodynamic rules, protein-protein interactions are responsible for a vast set of cellular functions. While some of them are composed of folded globular domains within stable complexes, others assemblies correspond to transient interactions involving short linear motifs, often found in disordered regions of the proteins. All these interactions are involved in multiple tasks, finely controlling gene expression, transports, biosynthesis and catabolism, post-translational modifications, signaling cascades and cell cycle. Deregulation of these macromolecular systems is subject to pathologies and cancer.Accumulated over the years, the increase in data obtained from various experimental screenings brings an unprecedented quantity of information, opening the doors of a contemporary stage of awareness. The thousands of atomic resolution protein structures, alone or in contact with an endogenous or exogenous partner, allows to observe the large variety of complexes and adoptable conformations. While all this information become barely reachable to a human understanding, statistical analysis can decipher and learn from it. This authorizes now to reach an other level of potential, with accessible data enabling powerful machine-learning approaches now achieving sure and robust prediction at the multiple scales of biological system. Current advances in supervised learning amenable computer to extract the substance of the matter out of millions of experiments, virtual or not, and to assimilate the numerous laws regulating biological functions.In this manuscript are described novel tools for the description, modeling and predictions the interaction of either small chemical entities or short protein segments with structured domains. On one side, it focus on a set of 486 human protein-kinases which are responsible of regulating multiple cellular process and currently one of the major therapeutic drug targets in cancer, but for which only a subset are precisely characterized. ProteoChemoMetric machine-learning models were set up for the prediction of binding poses and binding affinities between small compounds, now enabling the profiling of their binders and yield guides for medicinal chemists. On the other side, the mining of protein-protein interaction databases throughout an intelligible and interactive interface allows to bring back the knowledge from large databases in a reachable manner to scientists.Both advances enable better computational predictions of the interacting partners. The above tools and approaches shall more confidently guide the design of compounds with desired physico-chemical properties, capable of reaching specific target and fine tuning and adjusting of biological functions.Hopefully these accomplishment will provide new insights and directions for the discovery of improved therapeutic solutions for the treatments of diseases.