EmoWOZ: A Large-Scale Corpus and Labelling Scheme for Emotion Recognition in Task-Oriented Dialogue Systems
Authors/Creators
- 1. Heinrich Heine University Düsseldorf
Description
This is the dataset created for the paper, "EmoWOZ: A Large-Scale Corpus and Labelling Scheme for Emotion Recognition in Task-Oriented Dialogue Systems" (https://arxiv.org/abs/2109.04919).
EmoWOZ is based on MultiWOZ, a multi-domain task-oriented dialogue dataset (https://github.com/budzianowski/multiwoz). It contains more than 11K task-oriented dialogues with more than 83K emotion annotations of user utterances. In addition to Wizard-of-Oz dialogues from MultiWOZ, we collect human-machine dialogues within the same set of domains to sufficiently cover the space of various emotions that can happen during the lifetime of a data-driven dialogue system. There are 7 emotion labels, which are adapted from the OCC emotion models.
For data format and label definition, please refer to README.md.
A dataset sample containing 100 dialogues is released first. The full dataset will come soon.
Files
README.md
Files
(1.9 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:cf9aa8f7c26b7c4de1b6864a6aed1bdc
|
1.3 kB | Preview Download |
|
md5:6ececf3e20d254a882a3dcec0e02773a
|
1.9 MB | Preview Download |
Additional details
Related works
- Is published in
- Dataset: arXiv:2109.04919 (arXiv)