Published March 31, 2022 | Version v1

Oral cancer speech corpus for the paper "Objective speech outcomes after surgical treatment for oral cancer: An acoustic analysis of a spontaneous speech corpus containing 32.850 tokens"

  • 1. University of Groningen
  • 2. Netherlands Cancer Institute
  • 3. Netherlands Caner Institute

Description

Dataset accompanying the paper "Objective speech outcomes after surgical treatment for oral cancer: An acoustic analysis of a spontaneous speech corpus containing 32.850 tokens"

The zip file contains five folders:

- Database: contains csv files for each speaker which contain the processed features

- Recordings: the original recording from the YouTube Oral Cancer speech dataset, without further preprocessing

- Recordings_Normalised: same as recordings but after minimal audio preprocessing (min-max scaling)

- Textgrids: contains the textgrids which are annotated on the word-level and on phoneme-level

- TIMIT selection: contains the textgrids for the TIMIT speakers. We unfortunately cannot share the audio date as it is not open source. More information can be found here.

Files

dataset_oral_cancer_phoneme.zip

Files (603.2 MB)

Name Size
md5:ba4793e35de508e6dddbd3bf450d7497
603.2 MB Preview Download

Additional details

Funding

European Commission
TAPAS - Training Network on Automatic Processing of PAthological Speech 766287