Published June 8, 2021 | Version v1

CrowdSpeech and Vox DIY: Benchmark Dataset for Crowdsourced Audio Transcription

  • 1. Yandex
  • 2. Carnegie-Mellon University

Description

We collect and release CrowdSpeech — the first publicly available large-scale dataset of crowdsourced audio transcriptions. e show its applicability on an under-resourced language by constructing VoxDIY — a counterpart of CrowdSpeech for the Russian language.

Files

crowdspeech-dev-clean-gt.txt

Files (44.4 MB)

Name Size Download all
md5:e1891905ac533e2b46e4c503e142b575
3.5 MB Download
md5:1e3f2e61677e2f8047e0da89f7db72b0
490.1 kB Preview Download
md5:f7eca7edd3c18587354b10121f52e591
3.4 MB Download
md5:e61e89f0842f7979b46dec4f3741e4cf
479.6 kB Preview Download
md5:c8a83fd0cb6b158f51baf6ed1cb310c3
3.4 MB Download
md5:11184e82f9ddbb6eeb47aca0c57638ac
479.5 kB Preview Download
md5:acc5cc260329036e0797ee6aa325518e
3.5 MB Download
md5:78bf9a414a73682648427f8b1ead1156
495.0 kB Preview Download
md5:507bbe7c1b7058a997425605a4cca647
19.9 MB Download
md5:2d2316ebfe9b1d0f7971e4c697e507da
2.9 MB Preview Download
md5:8f3945c778aafa168d7d89e897172473
5.3 MB Download
md5:192e203ff41899a25f563615dc9177c7
746.5 kB Preview Download

Additional details

Related works

Is supplement to
Conference paper: http://openreview.net/forum?id=3_hgF1NAXU7 (URL)
Conference paper: https://arxiv.org/abs/2107.01091 (URL)