Logo image
Resources to manage Rice Big Genomics Data
Poster de colloque

Resources to manage Rice Big Genomics Data

Clément Agret, Celine Gottin, Alexis Dereeper, Christine Tranchant-Dubreuil, Annie Chateau, Anne Diévart, Gautier Sarah, Alban Mancheron, Manuel Ruiz et Gaëtan Droc
International Symposium of Rice Functional Genomics (Taipei, Taiwan, 04/11/2019–06/11/2019)
2019

Résumé

Constant progress in sequencing technologies creates a huge data overload. These big data combine Volume, Velocity and Variability constraints. Therefore, it is crucial to be able to integrate large amounts of heterogeneous data, with different formats and semantics, and manipulate them through complex workflows. This requires new, automated methods and tools for data integration and workflow management, to enable users with different backgrounds and interests to easily integrate and manipulate various datasets. We have developed the Rice Genome Hub, an integrative genome information system that allows centralized access to genomics and genetics data, and analytical tools to facilitate translational and applied research in rice. The hub is built using the Content Management System Drupal with the Tripal module that interacts with the Chado database. The Hub interface provides several functionalities (Blast, DotPlots, Gene Search, JBrowse, Primer Blaster, Primer Designer) to make it easy for querying, visualizing and downloading research data. We also plugged in-house tools developed by the South Green bioinformatics platform such as SNiPlay (detection and analyses of SNPs), Gigwa (filtering on genomic variations), daTALbase (exploration of data related to Xanthomonas TAL effectors), and DiffExDB (differential expression analysis).We also developed RedOak, a reference-free and alignment-free software package that allows for the indexing of a large collection of similar genomes. RedOak can be applied to reads from unassembled genomes, and it provides a nucleotide sequence query function. This software is based on a k-mer approach and has been developed to be heavily parallelized and distributed on several nodes of a cluster. Analysis of presence-absence variation (PAV) of genes among different genomes is a classical output of pan-genomic approaches. RedOak has a nucleotide sequence query function, including reverse complements, that can be used to quickly analyze the PAV of a specific gene among a large collection of genomes.

Fichiers et liens (1)

url
Find in HALAfficher

Indicateurs

1 Consultations de la notice

Détails

Logo image