3709626
doi
10.5281/zenodo.3709626
oai:zenodo.org:3709626
user-newseye
user-eu
Doucet, Antoine
University of La Rochelle
Lejeune, Gael
Sorbonne University
Odeo, Moses
Multimedia University Kenya
A Dataset for Multi-lingual Epidemiological Event Extraction
Mutuvi, Stephen
Multimedia University Kenya
info:eu-repo/semantics/openAccess
Creative Commons Attribution 4.0 International
https://creativecommons.org/licenses/by/4.0/legalcode
Epidemiology
Corpus Creation
Event Extraction
Classification
Multingual
<p>This paper proposes a corpus for development and evaluation of tools and techniques for identifying emerging infectious disease threats in online news text. The corpus can not only be used for Information Extraction, but also for other Natural Language Processing tasks such as text classification. We make use of articles published on the Program for Monitoring Emerging Diseases (PROMED) platform, which provides current information about outbreaks of infectious disease globally. Among the key pieces of information present in the articles is the Uniform Resource Locator (URL) to the online news sources where the outbreaks were originally reported. We detail the procedure followed to build the dataset, which include leveraging the source URLs to retrieve the news reports and subsequently pre-processing the retrieved documents. We also report on experimental results of event extraction on the dataset using the Data Analysis for Information Extraction in any Language(DANIEL) system. DANIEL is a multilingual news surveillance system that leverages unique attributes associated with news reporting repetition and saliency, to extract events. The system has a wide geographical and language coverage, including low-resource languages. In addition, we compare different classification approaches in terms of their ability to differentiate between epidemic related and non-related news articles that constitute the corpus.</p>
<p><strong>Dataset</strong></p>
<p>In addition to the paper, you may also be interested in <a href="https://zenodo.org/record/3709617">the dataset</a>.</p>
Zenodo
2020-05-13
info:eu-repo/semantics/conferencePaper
3693646
user-newseye
user-eu
award_title=NewsEye: A Digital Investigator for Historical Newspapers; award_number=770299; award_identifiers_scheme=url; award_identifiers_identifier=https://cordis.europa.eu/projects/770299; funder_id=00k4n6c32; funder_name=European Commission;
award_title=Cross-Lingual Embeddings for Less-Represented Languages in European News Media; award_number=825153; award_identifiers_scheme=url; award_identifiers_identifier=https://cordis.europa.eu/projects/825153; funder_id=00k4n6c32; funder_name=European Commission;
1644415621.293818
132556
md5:a2bb68b2d8b7ddf9ae947caa32449704
https://zenodo.org/records/3709626/files/A_Dataset_for_Epidemiological_Event_Extraction__CORRECTIONS(1).pdf
public
10.5281/zenodo.3693646
isVersionOf
doi