Published July 27, 2026 | Version 0.2

xSID-audio

  • 1. ROR icon Ludwig-Maximilians-Universität München
  • 2. ROR icon Munich Center for Machine Learning

Description

xSID-audio 0.2

xSID-audio consists of audio recordings for xSID, a multilingual dataset for slot and intent recognition. For more details and a data statement, see:

Release 0.2 of xSID-audio contains audio recordings for the validation and test sets of the xSID 0.5 splits for German (de) and for Bavarian as spoken in rural Upper Bavaria (de-ba). xSID-audio consists of read speech. The audios for German and Bavarian are recorded by the same speaker. More details are in the data statement in the appendix of Blaschke, Winkler & Plank (2026).

Updates:

  • v0.2: Updates to the metadata TSV files (no updates to the audio files):
    • A. This fixes a misalignment between the Bavarian sentence IDs on the one hand, and the German sentences & Bavarian audios on the other hand in the xSID test files. This part of the update corresponds exactly to how we used the data in the experiments for the ACL 2026 paper.
    • B. It also provides updated text versions of a few xSID sentences whose original translations deliberately included orthographical/grammatical errors (see Appendix F of van der Goot et al., 2021), which are fixed in the audio version to make the recordings more natural, or abbreviations which were expanded for the audio recordings (e.g., "AK" -> "Alaska", "Nov." -> "November"). The Text column contains the recorded version, Text_original the version as it appears in xSID.
  • v0.1: Initial release.

Using and citing xSID-audio

We share the audio recordings for research on processing spoken language data, but do not permit their use in the context of speech synthesis or voice cloning.

If you use xSID-audio, please cite the publication in which we introduce the dataset:

Please also cite the publications associated with xSID listed at the bottom of this file (at least the publications for related to the German and Bavarian xSID data: van der Goot et al., 2021, and Winkler et al., 2024).

Folder structure and file types:

  • The folders de and de-ba contain the recordings for the German and Bavarian data, respectively.
  • They each contain a subfolder for the validation (valid) and test (test) splits.
  • Each recording is a WAV file (recording rate: 48 kHZ, bit rate: 768 kBit/s).
  • The TSV files in the top-level folder (xsid_{de,de-ba}_{valid,test}.tsv) contain the instance ID for each sentence, the text from xSID 0.7, the intent from xSID 0.7, and the corresponding audio path.

Source datasets

xSID was published by van der Goot et al. (2021) [v0.1], and extended by Aepli et al. (2023) [v0.4], Winkler et al. (2024) [v0.5], Mæhlum & Scherrer (2024) [v0.7], and Krückl et al. (2025) [v0.7]. xSID is available at https://github.com/mainlp/xsid and licensed under a CC BY-SA 4.0 Intl. license. xSID builds on datasets by Coucke et al. (2018) – https://github.com/sonos/nlu-benchmark, license: CC 0 v.10 Universal –, and by Schuster et al. (2019) – https://fb.me/multilingual_task_oriented_data, license: CC BY-SA.

Files

de-ba.zip

Files (430.5 MB)

Name Size
md5:828a13112562562887add9b5160644e1
209.3 MB Preview Download
md5:10a2d23acec7773fd0dc29d484d0cd3b
221.0 MB Preview Download
md5:8b86781d2d273045875fb378d4ee2625
7.1 kB Preview Download
md5:2aaa58f5f6cf7a69aa47ed766deede33
70.2 kB Download
md5:b004d375d5b30e8640a005d4651a1acb
42.4 kB Download
md5:fbf954585cdaaa682927e6ac659fbd47
71.5 kB Download
md5:31e7ab6acbdd11f87c77839331efd8fc
43.7 kB Download

Additional details

Funding

European Commission
DIALECT - Natural Language Understanding for non-standard languages and dialects 101043235