There is a newer version of the record available.

Published March 4, 2024 | Version 1.0.0
Dataset Restricted

NeuroVoz: a Castillian Spanish corpus of parkinsonian speech

Description

The NeuroVoz dataset emerges as a pioneering resource in the field of computational linguistics and biomedical research, specifically designed to enhance the diagnosis and understanding of Parkinson's Disease (PD) through speech analysis. This dataset is distinguished as the first of its kind to be made publicly available in Castilian Spanish, addressing a critical gap in the availability of linguistic and dialectical diversity within PD research.

Compiled from a cohort of 108 participants, including 53 individuals diagnosed with PD and 55 healthy controls, the NeuroVoz dataset offers a rich compilation of speech recordings. All PD participants were recorded under medication (ON state), ensuring consistency and reliability in the speech samples collected. The dataset is meticulously curated to include a variety of speech tasks—ranging from sustained vowel phonations and diadochokinetic (DDK) tests to 16 structured listen-and-repeat utterances and spontaneous monologues. The inclusion of both manually transcribed listen-and-repeat tasks and Whisper-automated transcriptions for monologues underscores our commitment to data accuracy and usability.

Encompassing 2,903 audio files, the NeuroVoz dataset provides an extensive repository, averaging 26.88+- 3.35 recordings per participant, making it an invaluable asset for researchers seeking to explore the nuances of PD-affected speech. The dataset's structure and composition facilitate a multifaceted analysis of speech impairments associated with PD, offering insights into phonatory, articulatory, and prosodic changes.

In contributing to the body of knowledge with the NeuroVoz dataset, we invite the scientific community to engage with this dataset, explore the specific speech characteristics of PD in Castilian Spanish speakers, and advance the field of PD diagnosis through innovative speech analysis techniques.

 

If you use this dataset, please cite both this Zenodo and the arXiv preprint:

  • arXiv preprint: J. Mendes-Laureano, J. A. Gómez-García, A. Guerrero-López,E. Luque-Buzo, J. D. Arias-Londoño, F. J. Grandas-Pérez, and J. I. Godino-Llorente, “Neurovoz: a castillian spanish corpus of parkinsonian speech,” arXiv preprint arXiv:2403.02371 (2024).
    • Link: https://arxiv.org/abs/2403.02371
  • Zenodo dataset: Mendes-Laureano, J., Gómez-García, J. A., Guerrero-López, A., Luque-Buzo, E., Arias-Londoño, J. D., Grandas-Pérez, F. J., & Godino Llorente, J. I. (2024). NeuroVoz: a Castillian Spanish corpus of parkinsonian speech (1.0.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.10777657

Files

Restricted

The record is publicly accessible, but files are restricted to users with access.

Additional details

Funding

Agencia Estatal de Investigación
Ministry of Economy and Competitiveness of Spain PID2021-128469OB-I00
Agencia Estatal de Investigación
Ministry of Economy and Competitiveness of Spain TED2021-131688B-I00
Universidad Politécnica de Madrid
Maria Zambrano 2021 Maria Zambrano 2021
Agencia Estatal de Investigación
Ministry of Economy and Competitiveness of Spain DPI2017-83405-R1