Published June 12, 2023 | Version v2

RescueSpeech: A German Corpus for Speech Recognition in Search and Rescue Domain

  • 1. Saarland University, Germany
  • 2. German Research Center for Artificial Intelligence (DFKI), Saarbrucken

Description

Dear User,

We are thrilled to introduce our latest release - the RescueSpeech audio dataset, comprising authentic German speech recordings obtained from simulated search and rescue (SAR) exercises. The dataset contains manually annotated recordings from native German speakers, which were initially captured at 44.1 kHz and later down-sampled to 16 kHz to obtain a set of mono-speaker-single channel audio recordings. In order to protect the identity of the speakers, their names have been anonymized.

The RescueSpeech dataset is divided into two sets, each designed for different tasks: Automatic Speech Recognition (ASR) and Speech Enhancement.

1. For the ASR task, the dataset spans a duration of 1 hour and 36 minutes. It comprises a collection of clean-noisy pairs, where the noisy utterances are created by introducing contaminations from five different noise types sourced from the AudioSet dataset. These noise types include emergency vehicle siren, breathing, engine, chopper, and static radio noise. To match the 2412 clean utterances in the dataset, we have synthesized an equal number of corresponding noisy utterances. Additionally, we have provided the noise waveform files used to create the noisy utterances, ensuring transparency and reproducibility in the research community.

2. The Speech Enhancement task dataset is larger in size compared to the ASR dataset. The primary objective of this dataset is to facilitate the fine-tuning of speech enhancement models, particularly for the five SAR noise types mentioned earlier: emergency vehicle siren, breathing, engine, chopper, and static radio noise. Given the limited duration of clean audio available (1 hour and 36 minutes), we have synthesized multiple noisy utterances with varying noise types and signal-to-noise ratio (SNR) levels, all derived from a single clean utterance. This augmentation approach allows us to generate a more extensive dataset for speech enhancement purposes while preserving the original speaker distribution.

By providing these diverse datasets, we aim to support advancements in ASR and Speech Enhancement research, enabling the development and evaluation of robust systems that can handle real-world scenarios encountered during search and rescue operations.
 

Notes

Our work was supported under the project "A-DRZ: Setting up the German Rescue Robotics Center" and funded by the German Ministry of Education and Research (BMBF), grant No. I3N14856.

Files

Readme.md

Files (4.0 GB)

Name Size
md5:5b6bf9922c78e37aef96199f0eb34b4d
6.4 kB Preview Download
md5:c5b9f7857c1970bfe34eae0805c6987b
643.2 MB Download
md5:0a933817af7ccd56982225ae46c3722c
3.3 GB Download

Additional details

Related works

Is cited by
Conference paper: 10.48550/arXiv.2306.04054 (DOI)