Contrastively focused pronouns
Description
Data used in the Interspeech 2022 paper "BERT, can HE predict contrastive focus? Predicting and controlling prominence in neural TTS using a language model"
This is a corpus of literary texts that have been annotated with prominence and boundary features using the Wavelet Prosody Toolkit (https://github.com/asuni/wavelet_prosody_toolkit). Each text is read by three separate speakers. A subcorpus of contrastively focused pronouns is also provided.
Train and Test sets for prominence prediction task.
-Lines starting with '<file>' identify the utterances. Here you will find book/speaker/chapter/chapter-utterance# information.
-The columns of the remaining lines:
word / quantized CWT prominence features / quantized CWT boundary features / Raw CWT prominence features / Raw CWT boundary features
Majority, Minority and nonContrastivePronouns are dictionaries containing the chapter-utterance# (in Test.txt) and sentence index of the pronouns used for evaluation in the Interspeech paper.
Majority - at least two out of three speakers use contrastive focus.
Minority - only one speaker used contrastive focus.
nonContrastivePronouns - none of the speakers used contrastive focus.
Files
Majority.txt
Files
(66.1 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:b552c7b63edddc84349823f6cd0ac8ac
|
5.5 kB | Preview Download |
|
md5:9735fcbfaa6617e856d1c9f7f4080adf
|
1.5 kB | Preview Download |
|
md5:a81aac88d2be7d651675c27711ca96fb
|
6.1 kB | Preview Download |
|
md5:83606313a92b50766a8dd89170b61d11
|
649 Bytes | Preview Download |
|
md5:0206c9bf4521fd7485081f2facfb6674
|
9.8 MB | Preview Download |
|
md5:86101ed6d999ea0e34b3e040dd857bdb
|
56.3 MB | Preview Download |