Published February 3, 2025 | Version v1

Fused dataset

  • 1. ROR icon Brunel University of London

Description

The current version of the TOD (Task Oriented Dialogues) fused dataset contains samples from MultiWOZ2.2 (Zang et al., 2020), SpokenWOZ (Si et al., 2023), FRAMES (Asri et al., 2017), DSTC3 (Henderson et al., 2014a) and SGD (Rastogi et al., 2020) datasets. These datasets have been selected due to them all being high quality, with significant human validation and data cleaning. Additionally, this selection of datasets provides coverage across unique attributes, such as utterance-level audio files (Si et al., 2023).

The fused dataset requires several domains, necessitated by the scope of ELOQUENCE project (https://eloquenceai.eu) and the individual pilots. These datasets are stored using the ‘.arrow’ file extension so that speed and efficiency of data loading is optimised, as well as being compliant with the popular HuggingFace dataset library (HuggingFace, 2024). The dataset is also available at https://huggingface.co/datasets/Brunel-AI/ELOQUENCE. Currently, several datasets have been implemented within this fused dataset. However, due to the flexibility with which the schema has been defined, there is scope for additional datasets to be implemented across later iterations as further needs are identified. The JSON schema, as well as further explanation for attributes across all domains, is provided within Appendix 10.1 in ELOQUENCE deliverable 1.1.

Files

fused_dataset v1.0.zip

Files (65.8 MB)

Name Size Download all
md5:8322e687b39bb73ad928b8bcf360234e
65.8 MB Preview Download

Additional details

Funding

European Commission
ELOQUENCE - Multilingual and Cross-cultural interactions for context-aware, and bias-controlled dialogue systems for safety-critical applications 101135916

Software

Repository URL
https://huggingface.co/datasets/Brunel-AI/ELOQUENCE
Programming language
Python
Development Status
Active