Abstract
The extraction of significant biological patterns, and in particular the identification of regulation sites of proteinic synthesis in DNA primary sequences, is one of the major issues today in bioinformatics. Indeed any anomaly in proteinic synthesis regulation has detrimental damages on the well-being of certain organisms. Extracting these sites enables to better understand cellular operation or even to remove or cure pathology.<br /><br />What is problematic is the lack of information on patterns to be extracted, as well as the large volume of data to mine. In this dissertation, we introduce two polynomial algorithms -- the first one is deterministic and the other one is probabilist -- to address the issue of pattern extraction. We introduce a new family of score functions and we study their statistical properties. We characterize the language which is recognized by the index structure named "Oracle", and we modify this structure in order to make it more efficient.