Published June 15, 2022 | Version 1.0

Contrastively focused pronouns

Authors/Creators

  • 1. Gipsa-lab

Description

Data used in the Interspeech 2022 paper "BERT, can HE predict contrastive focus? Predicting and controlling prominence in neural TTS using a language model"

This is a corpus of literary texts that have been annotated with prominence and boundary features using the Wavelet Prosody Toolkit (https://github.com/asuni/wavelet_prosody_toolkit). Each text is read by three separate speakers. A subcorpus of contrastively focused pronouns is also provided.

Train and Test sets for prominence prediction task.

-Lines starting with '<file>' identify the utterances. Here you will find book/speaker/chapter/chapter-utterance# information.

-The columns of the remaining lines:

word  /  quantized CWT prominence features / quantized CWT boundary features  /   Raw CWT prominence features  /   Raw CWT boundary features

Majority, Minority and nonContrastivePronouns are dictionaries containing the chapter-utterance# (in Test.txt) and sentence index of the pronouns used for evaluation in the Interspeech paper.

Majority - at least two out of three speakers use contrastive focus.

Minority - only one speaker used contrastive focus.

nonContrastivePronouns - none of the speakers used contrastive focus.

Files

Majority.txt

Files (66.1 MB)

Name Size Download all
md5:b552c7b63edddc84349823f6cd0ac8ac
5.5 kB Preview Download
md5:9735fcbfaa6617e856d1c9f7f4080adf
1.5 kB Preview Download
md5:a81aac88d2be7d651675c27711ca96fb
6.1 kB Preview Download
md5:83606313a92b50766a8dd89170b61d11
649 Bytes Preview Download
md5:0206c9bf4521fd7485081f2facfb6674
9.8 MB Preview Download
md5:86101ed6d999ea0e34b3e040dd857bdb
56.3 MB Preview Download