xSID-audio
Authors/Creators
Description
xSID-audio 0.2
xSID-audio consists of audio recordings for xSID, a multilingual dataset for slot and intent recognition. For more details and a data statement, see:
- Verena Blaschke, Miriam Winkler, and Barbara Plank. 2026. Standard-to-Dialect Transfer Trends Differ across Text and Speech: A Case Study on Intent and Topic Classification in German Dialects. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 6807–6829, San Diego, California, United States. Association for Computational Linguistics.
Release 0.2 of xSID-audio contains audio recordings for the validation and test sets of the xSID 0.5 splits for German (de) and for Bavarian as spoken in rural Upper Bavaria (de-ba). xSID-audio consists of read speech. The audios for German and Bavarian are recorded by the same speaker. More details are in the data statement in the appendix of Blaschke, Winkler & Plank (2026).
Updates:
- v0.2: Updates to the metadata TSV files (no updates to the audio files):
- A. This fixes a misalignment between the Bavarian sentence IDs on the one hand, and the German sentences & Bavarian audios on the other hand in the xSID test files. This part of the update corresponds exactly to how we used the data in the experiments for the ACL 2026 paper.
- B. It also provides updated text versions of a few xSID sentences whose original translations deliberately included orthographical/grammatical errors (see Appendix F of van der Goot et al., 2021), which are fixed in the audio version to make the recordings more natural, or abbreviations which were expanded for the audio recordings (e.g., "AK" -> "Alaska", "Nov." -> "November"). The
Textcolumn contains the recorded version,Text_originalthe version as it appears in xSID.
- v0.1: Initial release.
Using and citing xSID-audio
We share the audio recordings for research on processing spoken language data, but do not permit their use in the context of speech synthesis or voice cloning.
If you use xSID-audio, please cite the publication in which we introduce the dataset:
- Verena Blaschke, Miriam Winkler, and Barbara Plank. 2026. Standard-to-Dialect Transfer Trends Differ across Text and Speech: A Case Study on Intent and Topic Classification in German Dialects. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 6807–6829, San Diego, California, United States. Association for Computational Linguistics.
Please also cite the publications associated with xSID listed at the bottom of this file (at least the publications for related to the German and Bavarian xSID data: van der Goot et al., 2021, and Winkler et al., 2024).
Folder structure and file types:
- The folders
deandde-bacontain the recordings for the German and Bavarian data, respectively. - They each contain a subfolder for the validation (
valid) and test (test) splits. - Each recording is a WAV file (recording rate: 48 kHZ, bit rate: 768 kBit/s).
- The TSV files in the top-level folder (
xsid_{de,de-ba}_{valid,test}.tsv) contain the instance ID for each sentence, the text from xSID 0.7, the intent from xSID 0.7, and the corresponding audio path.
Source datasets
xSID was published by van der Goot et al. (2021) [v0.1], and extended by Aepli et al. (2023) [v0.4], Winkler et al. (2024) [v0.5], Mæhlum & Scherrer (2024) [v0.7], and Krückl et al. (2025) [v0.7]. xSID is available at https://github.com/mainlp/xsid and licensed under a CC BY-SA 4.0 Intl. license. xSID builds on datasets by Coucke et al. (2018) – https://github.com/sonos/nlu-benchmark, license: CC 0 v.10 Universal –, and by Schuster et al. (2019) – https://fb.me/multilingual_task_oriented_data, license: CC BY-SA.
- Rob van der Goot, Ibrahim Sharaf, Aizhan Imankulova, Ahmet Üstün, Marija Stepanović, Alan Ramponi, Siti Oryza Khairunnisa, Mamoru Komachi, and Barbara Plank. 2021. From Masked Language Modeling to Translation: Non-English Auxiliary Tasks Improve Zero-shot Spoken Language Understanding. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2479–2497, Online. Association for Computational Linguistics.
- Noëmi Aepli, Çağrı Çöltekin, Rob Van Der Goot, Tommi Jauhiainen, Mourhaf Kazzaz, Nikola Ljubešić, Kai North, Barbara Plank, Yves Scherrer, and Marcos Zampieri. 2023. Findings of the VarDial Evaluation Campaign 2023. In Tenth Workshop on NLP for Similar Languages, Varieties and Dialects (VarDial 2023), pages 251–261, Dubrovnik, Croatia. Association for Computational Linguistics.
- Miriam Winkler, Virginija Juozapaityte, Rob van der Goot, and Barbara Plank. 2024. Slot and Intent Detection Resources for Bavarian and Lithuanian: Assessing Translations vs Natural Queries to Digital Assistants. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), pages 14898–14915, Torino, Italia. ELRA and ICCL.
- Petter Mæhlum and Yves Scherrer. 2024. NoMusic - The Norwegian Multi-Dialectal Slot and Intent Detection Corpus. In Proceedings of the Eleventh Workshop on NLP for Similar Languages, Varieties, and Dialects (VarDial 2024), pages 107–116, Mexico City, Mexico. Association for Computational Linguistics.
- Xaver Maria Krückl, Verena Blaschke, and Barbara Plank. 2025. Improving Dialectal Slot and Intent Detection with Auxiliary Tasks: A Multi-Dialectal Bavarian Case Study. In Proceedings of the 12th Workshop on NLP for Similar Languages, Varieties and Dialects, pages 128–146, Abu Dhabi, UAE. Association for Computational Linguistics.
- Alice Coucke, Alaa Saade, Adrien Ball, Théodore Bluche, Alexandre Caulier, David Leroy, Clément Doumouro, Thibault Gisselbrecht, Francesco Caltagirone, Thibaut Lavril, Maël Primet, and Joseph Dureau. 2018. Snips Voice Platform: an embedded Spoken Language Understanding system for private-by-design voice interfaces. Preprint on arXiv: 1805.10190.
- Sebastian Schuster, Sonal Gupta, Rushin Shah, and Mike Lewis. 2019. Cross-lingual Transfer Learning for Multilingual Task Oriented Dialog. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 3795–3805, Minneapolis, Minnesota. Association for Computational Linguistics.
Files
de-ba.zip
Files
(430.5 MB)
| Name | Size | |
|---|---|---|
|
md5:828a13112562562887add9b5160644e1
|
209.3 MB | Preview Download |
|
md5:10a2d23acec7773fd0dc29d484d0cd3b
|
221.0 MB | Preview Download |
|
md5:8b86781d2d273045875fb378d4ee2625
|
7.1 kB | Preview Download |
|
md5:2aaa58f5f6cf7a69aa47ed766deede33
|
70.2 kB | Download |
|
md5:b004d375d5b30e8640a005d4651a1acb
|
42.4 kB | Download |
|
md5:fbf954585cdaaa682927e6ac659fbd47
|
71.5 kB | Download |
|
md5:31e7ab6acbdd11f87c77839331efd8fc
|
43.7 kB | Download |
Additional details
Funding
Software
- Repository URL
- https://github.com/mainlp/dialects-text-vs-speech