Published May 21, 2026 | Version V0.1

HEREDITermCorpus_it

  • 1. Centro de Linguística da Universidade NOVA de Lisboa
  • 2. ROR icon Universidade Nova de Lisboa

Description

In the context of the HEREDITARY project, HetERogeneous sEmantic Data integratIon for the guT-brAin interplay, dedicated multilingual corpora are being created. The HEREDITermCorpus_it_V0.1 compiles a curated selection of texts dedicated to the microbiota-gut-brain axis (MGBA) and its emerging role in neurodegenerative disorders. The collection is intended to provide a resource for researchers, clinicians, and students interested in exploring how intestinal microorganisms influence brain health and disease mechanisms. The dataset comprises 206 documents, 144,009 sentences, 3,112,322 words and 4,140,747 tokens. All documents are written in Italian and were selected to capture a wide range of perspectives on the MGBA.

Files

Files (98.7 kB)

Name Size Download all
md5:0a512b8fafdf63254a2a95af0a827014
98.7 kB Download

Additional details

Related works

Is supplement to
Dataset: 10.5281/zenodo.16968962 (DOI)
Dataset: 10.5281/zenodo.16969241 (DOI)

Funding

European Commission
HEREDITARY - HetERogeneous sEmantic Data integratIon for the guT-bRain interplaY 101137074

Dates

Created
2026-03-05
Creation of the HEREDITermCorpus, for Italian texts
Available
2026-05-21
Version 0.1 of the HEREDITermCorpus, for Italian texts