Abstract
The molecular mechanisms governing the synthesis of proteins have historically received the most attention. However, the number of regions encoding proteins (coding regions) is limited. Mammalian genomes contain many non-coding regions that correspond to open chromatin and cis-regulatory modules bound by transcription factors (TFs). Another particularity of these genomes is that they are scattered with repetitive sequences, in particular microsatellites, also called short tandem repeats (STRs), which correspond to repeated DNA motifs of 1 to 6 bp and constitute one of the most polymorphic and abundant repetitive elements (~3% of the human genome). STRs are known to widely impact gene expression and contribute to expression variation. At the molecular level, STRs can for instance affect expression by inducing inhibitory DNA structures and/or by modulating TF binding. Specific expression Quantitative Trait Loci (eQTLs), called eSTRs, have been calculated to correlate STR length variations with gene expression.First, leveraging the Cap Analysis of Gene Expression (CAGE) data collected by the international FANTOM consortium, I contributed to the discovery of widespread transcription initiation at STRs in mouse and human cells. Interestingly, genetic variants linked to human diseases are located around STRs associated with high transcription initiation levels, suggesting a functional role of this transcription. I also demonstrated that this transcription initiation is predictable by sequence-based deep learning models. However, although methods exist to interpret this type of model after learning, CNN interpretation still has some limitations, impeding the acquisition of novel biological knowledge.To tackle this problem, I developed, in a second work, fully interpretable modular neural networks (MNNs), which combines learning and interpretation in one single step preserving automatic feature extraction from DNA sequence. I further used the predictions of these models to evaluate the impact of transcription initiation at STRs on gene expression and to compute a new catalog of eSTRs. Finally, I evaluated the phenotypic consequences of STR transcription using the data collected by the FINNGEN consortium which includes 224,737 genotypes and phenotypes.Overall, my work should serve as a valuable resource for future genetic studies of complex traits.