Abstract
Identifying the different parts of a biological sequence (nucleic sequence, or amino acid sequence) is a first step toward understanding the biology of the organism from which it originates. Given a set of biological sequences of an organism, we are interested in this thesis to the discovery of «domains», ie of relatively large subsequences (several tens of nucleotides or amino acids) that the we can find in a large number of sequences. This thesis is decomposed into two parts corresponding to the discovery of domains in the protein sequences and in the nucleic sequences. In each part, the methods developed are applied to Plasmodium falciparum, the pathogen responsible for malaria in humans, and for which conventional bioinformatic methods struggle to produce satisfactory annotations. The first developed part relates to the discovery of domains in protein sequences. A common approach to identifying domains of a protein is to perform sequencesequence comparisons with local alignment tools such as BLAST. However, these approaches sometimes lack sensitivity, particularly for species phylogenetically distant from conventional reference organisms. Here we propose an approach to increase the sensitivity of sequence-sequence comparisons. This new approach uses the fact that protein domains tend to appear with a limited number of other domains on the same protein. In Plasmodium falciparum, this method allows the discovery of 2 240 new domains for which, in the majority of cases, there is no similar model in domain databases. The second developed part relates to the discovery of domains in regulation sequences (DNA sequences). Several studies have shown that there is a strong link between the nucleotide composition of particular regions (promoter sequences in particular) and the expression of genes. We propose here a new approach to discover automatically these regions, which we call regulation domains. More specifically, our approach is based on a strategy of iterative exploration of nucleotide compositions, from the simplest (dinucleotides) to the most complex (k-mers), as well as a supervised segmentation strategy to discover compositions and regions of interest. Using the domains thus identified, we show that the expression of Plasmodium falciparum genes can be predicted with good precision. Applied to various other eukaryotic species, this approach shows very different results depending on the species (between 40 and 70% correlation) which suggests a regulation mechanism probably shared by all eukaryotic species but whose importance varies from one species to another.