There is a newer version of the record available.

Published September 10, 2021 | Version 1.0

FONA corpus: Food & Nutrition Abstracts Multilingual corpus

  • 1. Barcelona Supercomputing Center

Description

The FONA corpus is a collection of case reports specifically selected to foster the development of Language Technologies, Text Mining and NLP for applications in the domain of food, nutrition and agriculture.

 

It contains a large collection of documents (titles and abstracts) with metadata information on their PMC availability, MeSH terms and language. In addition, a subset of the collection contains automatically recognized entities of the following categories:

  • medical procedures
  • symptoms
  • diseases
  • medications
  • occupational and demographic information
  • species (pathogens)
  • cancer morphology

 

The FONA corpus will be disclosed on September 21st at the IberHeLT workshop: https://sites.google.com/campus.ul.pt/iberhelt2021/home

Notes

Funded by the Plan de Impulso de las Tecnologías del Lenguaje (Plan TL).

Files

FONA-corpus.txt

Files (5 Bytes)

Name Size Download all
md5:d8e8fca2dc0f896fd7cb4cb0031ba249
5 Bytes Preview Download