There is a newer version of the record available.

Published January 17, 2022 | Version 1.0.0-alpha

EmoWOZ: A Large-Scale Corpus and Labelling Scheme for Emotion Recognition in Task-Oriented Dialogue Systems

Description

This is the dataset created for the paper, "EmoWOZ: A Large-Scale Corpus and Labelling Scheme for Emotion Recognition in Task-Oriented Dialogue Systems" (https://arxiv.org/abs/2109.04919).

EmoWOZ is based on MultiWOZ, a multi-domain task-oriented dialogue dataset (https://github.com/budzianowski/multiwoz). It contains more than 11K task-oriented dialogues with more than 83K emotion annotations of user utterances. In addition to Wizard-of-Oz dialogues from MultiWOZ, we collect human-machine dialogues within the same set of domains to sufficiently cover the space of various emotions that can happen during the lifetime of a data-driven dialogue system. There are 7 emotion labels, which are adapted from the OCC emotion models.

For data format and label definition, please refer to README.md. 

A dataset sample containing 100 dialogues is released first. The full dataset will come soon.

Files

README.md

Files (1.9 MB)

Name Size Download all
md5:cf9aa8f7c26b7c4de1b6864a6aed1bdc
1.3 kB Preview Download
md5:6ececf3e20d254a882a3dcec0e02773a
1.9 MB Preview Download

Additional details

Related works

Is published in
Dataset: arXiv:2109.04919 (arXiv)