Logo image
PADI-web corpus: Labeled textual data in animal health domain
Jeu de données   Open Access   Avec comité de lecture

PADI-web corpus: Labeled textual data in animal health domain

Julien Rabatel, Elena Arsevska et Mathieu Roche
Data in Brief, Vol.22, p.643-646
2019
PMID: 30671512

Résumé

Monitoring animal health worldwide, especially the early detec-tion of outbreaks of emerging pathogens, is one of the means ofpreventing the introduction of infectious diseases in countries(Collier et al., 2008)[3]. In this context, we developed PADI-web, aPlatform for Automated extraction of animal Disease Informationfrom the Web (Arsevska et al., 2016, 2018). PADI-web is a text-mining tool that automatically detects, categorizes and extractsdisease outbreak information from Web news articles. PADI-webcurrently monitors the Web forfive emerging animal infectiousdiseases, i.e., African swine fever, avian influenza including highlypathogenic and low pathogenic avian influenza, foot-and-mouthdisease, bluetongue, and Schmallenberg virus infection. PADI-webcollects Web news articles in near-real time through RSS feeds.Currently, PADI-web collects disease information from GoogleNews because of its international and multiple language coverage.We implemented machine learning techniques to identify therelevant disease information in texts (i.e., location and date of anoutbreak, affected hosts, their numbers and clinical signs). In orderto train the model for Information Extraction (IE) from newsarticles, a corpus in English has been manually labeled by domainexperts. This labeled corpus (Rabatel et al., 2017) is presented inthis data paper.

Fichiers et liens (2)

url
Find in HALAfficher
url
https://doi.org/10.1016/j.dib.2018.12.063Afficher
Publié (version de la notice) Ouvrir

Indicateurs

1 Consultations de la notice

Détails

Logo image