Résumé
Miropeat analysis of E. cuniculi chromosomes. Miropeat analysis was performed on various set of sequences. The three-by-three analysis was performed to extract coordinates of “r” sequence blocks. A. Comparative analysis of the 11 E. cuniculi chromosomes. Most repeated sequences are associated to chromosome extremities. B. Three-by-three analysis confirming the existence of sequence block. Left: comparison of chromosome I, IV and VIII enables the identification of r01, r02, r03, r04 repeats. The r15 sequence was found by comparison of chromosome I with itself. Part of r02, r03, R04 and r15 repeats will composed the EXT1 sequence block. Figure S2. Miropeat – EXT correspondence. A. Superimposed distributions of repeated elements detected with Miropeat software (boxes r01 to r02), EXT blocks (coloured arrows 1 to 10). The r01 repeat includes one rDNA unit (red arrow) and r04 is ascribed to dhfr-ts (dihydofolate reductase - thymidylate synthase) gene cluster. Five DNA segments are of unique type (be1 to be5). Some EXT blocks may consist in the clustering of “r” and “be” sequences. EXT8 was created on the basis of BLAST homologies. B. E. cuniculi chromosome extremities have been arranged by referring to the conserved position of R01 recombination site. The R02 recombination is present at chromosome ends presenting an S-to-EXT1 or S-to-EXT5 transition. Hachured boxes correspond to missing regions in the Genoscope E. cuniculi genome release that have been reassembled in our study. EXT9 and EXT10 were completely characterized after specific cloning and sequencing of IIIβ and IXα ends, respectively. Highly conserved regions at all chromosome ends (telomeres and distal subtelomeric regions) are symbolized by a red dashed arrow on both sides of the schema. The “r” and EXT repeats are represented at their correct scale. The scale was not respected for SUB and coding core regions because of graphical reasons. Figure S3. Experimental validation of the mosaic structure described for E. cuniculi GB-M1 subtelomeres. Long-PCR amplifications were used to reconstruct chromosome ends. Three hypotheses for linking sequences were tested and are illustrated here: from EXT block to chromosome core (e.g. Lr12dir2/LKVβ), from rDNA to chromosome core (e.g. Lr02r1-2/LKIXα) and from rDNA to EXT block (e.g. Lr01-2/Lr07rev). A. general strategy based of PCR and long-PCR amplifications. B. Example of PCR amplification products. Agarose gel analysis of PCR products that are sometimes more than 15 kbp in size. The DNA size marker is Lambda DNA cut with EcoRI and HindIII. C. Relative positions of the primers, the R01 recombination site being taken as an origin. Primers sequences and orientations are given in Additional file 2: Table S1. Figure S4. Experimental validation of the two IIIβ extremities genetic organization in E. cuniculi GB_M1 strain. A. Schematic representation of IIIβ1 chromosome end organization deduced from PCR products and sequencing. The relative positions of the different primers were identified. Most primers were design from EXT repeats characterized in other chromosome extremities. Coordinates are given according to the current release of E. cuniculi chromosome III. B. PCR and long-PCR amplifications were used to reconstruct chromosome IIIβ extremities. Agarose gel analysis reveals PCR products that are sometimes more than 15 kbp in size. The DNA size marker is Lambda DNA cut with EcoRI and HindIII. Figure S5. Coding DNA sequences (CDSs) and recombination sites (R01 to R18) in EXT repeats distributed among the 22 chromosome ends (Σα, Σβ) of Encephalitozoon cuniculi. Each CDS is schematized by a large arrow showing the direction of transcription. New putative CDS deduced from E. cuniculi GB-M1 chromosome ends framework of reconstruction are represented by thick outlined shapes. To gain space, CDS names have been reduced to the last four digit numbers (e.g. 0040 on Iα and 0040 on IVα correspond to ECU01_0040 and ECU04_0040 in E. cuniculi genome databases, respectively). If available, Uniprot accession numbers for encoded hypothetical proteins are also given (Y103_ENCCU, Y110_ENCCU…). CDSs assigned to four multigene inter families (AE, B, C or D) are coloured in grey. Recombination hot spots, indicated by vertical dashed lines (“precise” sites) or hatched zones (“approximate” sites), determine the boundaries between EXT sub-blocks (e.g. EXT1-1, a sub-block of EXT1, is bounded by R02 and R03 sites). CDSs surrounded by dotted line are in the coding core sequence. CDSs with plain lines are EXT and SUB CDSs. Arrows without number are extrapolated from the present study. A. EXT1. B. EXT2. The thick horizontal bare indicate the region covered by corresponding Genbank accession number. C. EXT3. The thick horizontal bare indicate the region covered by corresponding EMBL/Genbank accession number. D. EXT4. E. EXT5. The thick horizontal bare indicate the region covered by corresponding Genbank accession number. F. EXT6. G. EXT7. Figure S6. Gradient complexity of EXT repeats from R01 recombination to EXT-to-core transition. Recombination have place in a tree-like structure respecting their relative distance to R01 recombination site. All EXT repeats and the 23 chromosome ends are represented in the schema. Recombination site fusion are represented by a discontinuous line. Duplication of R09 recombination site is indicated by a plain dark line. Interestingly, R08 and R13 sites were at the same distance from R01 when they get fused. Figure S7. Progressive or rapid GC % shift at EXT-to-core transition. EXT repeats are associated with high GC %. This observed shift is improved by the low GC % that is present at EXT-to-core transition (35 % GC in average). It was possible to measure the distance between the lowest GC % peaks at one chromosome extremity with the highest GC % optimum found in the adjacent EXT repeat (red line). We consider only the first EXT subsequence in that experiment. In fact, we observe that low GC % values were associated with some recombination sites. The slope of the GC % curve varies a lot between chromosomes. Four examples of the most extreme values are given in the panel. Figure S8. GeneFizz analysis of an E. cuniculi chromosome end (IIα) showing the relationship between open-closed transitions and recombination sites. A. An overview of a complete IIα organisation can be obtained by combination of different sequences (see bottom). CDSs are represented by boxes with arrows indicating their direction of transcription and are ordered from telomere (left) to chromosome core (right). Letters AE, B, C and D are applied to the members of the four multigenic families described in the present study. Name and Uniprot accession number are given for most CDSs. B. DNA of EXT repeats is getting melted at high temperature and offers an alternation of “open” and “closed” areas. Double stranded DNA (“closed” state) correspond to the null value. At the maximum, the DNA is fully melted, i.e. single stranded (“open” state). Corresponding colours are red for 72 °C, yellow for 73 °C, deep blue for 74 °C, light blue for 75 °C, violet curve without colour shading for 76 °C. The green line corresponds to the GC %. An increase of at least 3 °C is required for strand dissociation at some loci (*). Recombination sites, indicated by vertical dashed lines (“precise” sites) or hatched zones (“approximate” sites), determine the boundaries between EXT sub-sequences. Small vertical arrows at the bottom indicate the position that was taken into consideration to calculate the distance between open-to-closed transition and recombination site. C. Recombination sites R10 and R11 are associated with close regions. Corresponding colours, from 69 to 73 °C, are red, yellow, dark blue, light blue and violet. GC content curve is coloured in green. The R11 site is associated with a closed region in the proximal region. Figure S9. Description and putative structure of InterAE, InterB, InterC and InterD proteins. Protein description was recovered from Pfam database. We use one entry per multigene family product which was considered to be the most representative. We recovered both graphical representation of the protein and feature table. A schematic representation can be deduced from each Pfam description. We consider the distribution of positively charged amino acids (+) around the first transmembrane domain to assess the orientation of the InterC and InterD protein in the plasma membrane as they have no ER-targeting signal peptide. Figure S10. MicrosporidiaDB resources presenting genomic distribution of interAE, interB, interC and interD genes in Encephalitozoon cuniculi genome. Data were recovered of the GB-M1 strain that was used in the present study. They were also recovered for two other isolates EC1 and EC3. A. The interAE and interB genes were selected from the database based on their conserved Interpro domain IPR011667-UPF0329. Some truncated genes and pseudogenes were not detected. B. The interC genes were selected from the database based on their conserved Pfam domain DUF1686. B. The interD genes were selected from the database based on their conserved Interpro domain IPR019081-UPF0328. Some partial sequences were not detected. (PDF 1180 kb)