Published July 31, 2020 | Version v1

Youtube-Dataset for Language Identification in Speech Signals

  • 1. Fraunhofer IDMT

Description

Youtube-Dataset for Language Identification in Speech Signals

- for scientific use only, for questions contact: jakob.abesser@idmt.fraunhofer.de

Reference

In case you use this dataset for your research, please cite

Alexandra Draghici, Jakob Abeßer & Hanna Lukashevich: A Study on Spoken Language Identification
using Deep Neural Networks, Proceedings of the Audio Mostly Conference 2020

Dataset

The YouTube News Collection is a collection of videos from various
Youtube news channels. We gathered data from channels like BBC
news, France24, DW News, and Noticias Telemundo.

- 135664 npy files (numpy matrices exported from Python)
- each npy file includes a mel spectrogram (see below) of an audio file
- the subfolders "0" - "5" encode the language id:
  0 - English
  1 - French
  2 - German
  3 - Greek
  4 - Italian
  5 - Spanish

Audio Processing

- mono, sample rate 22.05 kHz
- mel spectrogram (librosa python package)
- windows size 512 samples
- hopsize 441 samples (20 ms)
- 129 mel bands
- file-level spectrogram are normalized to maximum of 1
 

Files

LanguageID_Youtube_6_Classes.zip

Files (47.8 GB)

Name Size
md5:89f2cf96b617d6f5032ef261aa743ab3
47.8 GB Preview Download
md5:fc337bfaa4901fbe7b05f5ec641a237e
1.1 kB Preview Download