Résumé
MGT/HighV-QUEST has been developed by IMGT® [1], the international ImMunoGeneTics Information System® to answer the problematic of the analysis of the antigen receptor data from Next-Generation Sequencing (NGS). The analysis of the expressed repertoires of antigen receptors - immunoglobulins (IG) or antibodies and T cell receptors - represents a crucial challenge for the study of the adaptive immune response in normal and disease-related situations. Currently the standardized analysis of IG and TR nucleotide sequences, based on the IMGT-ONTOLOGY concepts [2,3], is performed by IMGT/V-QUEST [4], the integrated IMGT® tool available online for 50 or less IG and TR rearranged sequences. IMGT/HighV-QUEST is the high-throughput version of IMGT/V-QUEST that allows users to analyse batches of more than 100,000 rearranged sequences of antigen receptors, IG or TR, in one run. IMGT/HighV-QUEST is a secure system and web portal. It requires user identification and all data transactions are performed over a secure web connection provided by HTTPS technologies. It provides a simple user interface accessed via a classical web browser which is familiar to IMGT/V-QUEST users. In particular, the functionalities are identical to those provided for 'Detailed View' and 'Excel files' of the online IMGT/V-QUEST tool. The results comprise a set of text files which include:1) Eleven files equivalent to the eleven sheets of the 'Excel files' whose content is detailed in IMGT/V-QUEST Documentation: (i) the 'Summary' file provides the synthesis of the analysis (the sequence functionality, the names of the closest variable (V), diversity (D), and joining (J) genes and alleles with identity percentage, framework (FR) and complementarity determining region (CDR) lengths, amino acid (AA) V-D-J or V-J JUNCTION, the description of insertions and deletions if any…), (ii) the 'IMGT-gapped-nt-sequences' file includes the nucleotide (nt) sequences of labels that have been gapped according to the IMGT unique numbering, (iii) the 'nt-sequences' file includes the nt sequences of all described labels (iv) the 'IMGT-gapped-AA-sequences' file the amino acid sequence that have been gapped according to the IMGT unique numbering [5], (v) the 'AA-sequences file' includes the AA sequence of labels without IMGT gaps, (vi) the 'Junction' file includes the results of IMGT/JunctionAnalysis [6], (vii) the 'V-REGION-mutation-table' file includes the list of mutations (nt mutation, AA change, AA class identity or change per FR and CDR, (viii) the 'V-REGION-nt-mutation-statistics' file includes the number of positions including IMGT gaps, the number of nt, the number of identical nt, the total number of mutations, the number of silent mutations, the number of nonsilent mutations, the number of transitions and the number of transversions per FR and CDR, (ix) the 'V-REGION-AA-mutation-statistics' file includes the AA positions including IMGT gaps, the number of AA, the number of identical AA, the total number of AA changes, the number of AA changes according to AAclassChangeType, and the number of AA class changes according to AAclassSimilarityDegree per FR and CDR, (x) the 'V-REGION-mutation-hot-spots' file indicates the localization of the hot spots motifs detected in the closest germline V-REGION with positions in FR and CDR, and (xi) the 'Parameters' file includes the date of the analysis, the IMGT/V-QUEST version, and the parameters used for the analysis.2) The 'Detailed view' for each analysed sequence that allows visualizing the individual results: (i) the result summary summarizes the main characteristics of the analysed sequence with the names of the closest V and J genes and alleles with their alignment score and the percentage of identity, the name of the closest ‘D-gene and allele determined by IMGT/JunctionAnalysis with the D-REGION reading frame, the FR and CDR lengths and the AA JUNCTION sequence, and if selected, (ii) the Alignment for V, D, J genes and alleles, (iii) the detailed analysis of the JUNCTION by IMGT/JunctionAnalysis, (iv) different displays of the V-REGION, (v) the analysis of the mutations and AA changes, (vi) the localization of the mutation hot spots, and (vii) the annotation by IMGT/Automat [7].One of the challenges of IMGT/HighV-QUEST is the management of many analyses of thousands of sequences and of their results. The analyses are managed locally and the corresponding jobs are dispatched to computational servers provided they have available resources. This requires the establishment of a stateful system which keeps trace and follows the status of the analyses even when users are not connected. Another strategy taken to enforce data security is the encryption of links sent by E-mail to users to let them deal with the analysis results. This encryption using RSA algorithm and the requirement of a user session, to follow the analysis status or to deal with the results, enforce the security aspect of IMGT/HighV-QUEST. The result files are saved on file systems that are protected by firewalls and interactions with these file systems are only possible by SSH (Secure Shell) connections. The user results are kept for a limited period of time on file systems (15 days) and are fully deleted afterwards. The current release of IMGT/HighV-QUEST is in test on internal IMGT® servers. It will then be accessible from the IMGT® Home page (http://www.imgt.org).