Résumé
Words are one of the grounds of European languages. Corpora written with these languages are normally describe by words. However, extracted information given by words is semantically poor. Actually, to take into account the complexity of European languages are really important. As a result, we propose in this thesis to feature the characteristic of European languages by using syntactic informations in order to discover new semantic knowledge from corpora. First, we present SELDE, a model of feature selection. This one is based on objects extracted from syntactic relations of a corpus. We experiment SELDE on textual classification tasks by proposing Ex- pLSA, an approach used to make a corpus expansion by using the SELDE features. The goal of ExpLSA is to combine the SELDE features with the statistic method LSA. The SELDE model gives relevant features but cannot be apply with all kinds of textual data. Thus, we propose different approaches adapted to specific textual data, called complex textual data. We experiment our approaches with noised data, bad written data, and data without syntactic informations. Finally, we propose the SELDEF model. It introduce the automatic validation of syntactic relations called induced. Two validation approaches are proposed : a Semantic-Vector-based approach and a Web Validation system. The Semantic Vectors approach is a Roget-based method which computes a syntactic relation as a vector. Web Validation uses a search engine to determine the relevance of a syntactic relation. Then, we propose approaches to combine both in order to rank induced syntactic relations. We experiment SELDEF in a conceptual classes building task. Obtained results confirm the quality of validation approaches and quality of built classes.