MaSS - Multilingual corpus of Sentence-aligned Spoken utterances

doi:10.5281/zenodo.3354711

Published July 30, 2019 | Version 1.0

Dataset Open

MaSS - Multilingual corpus of Sentence-aligned Spoken utterances

1. Université Grenoble Alpes

Abstract

The CMU Wilderness Multilingual Speech Dataset is a newly published multilingual speech dataset based on recorded readings of the New Testament. It provides data to build Automatic Speech Recognition (ASR) and Text-to-Speech (TTS) models for potentially 700 languages. However, the fact that the source content (the Bible), is the same for all the languages is not exploited to date. Therefore, this article proposes to add multilingual links between speech segments in different languages, and shares a large and clean dataset of 8,130 para-lel spoken utterances across 8 languages (56 language pairs).We name this corpus MaSS (Multilingual corpus of Sentence-aligned Spoken utterances). The covered languages (Basque, English, Finnish, French, Hungarian, Romanian, Russian and Spanish) allow researches on speech-to-speech alignment as well as on translation for syntactically divergent language pairs. The quality of the final corpus is attested by human evaluation performed on a corpus subset (100 utterances, 8 language pairs).

Paper | GitHub Repository containing the scripts needed to build the data set from scratch (if needed)

Project structure

This repository contains 8 Numpy files, one for each featured language, pickled with Python 3.6. Each line corresponds to the spectrogram of the file mentioned in the file verses.csv. There is a direct mapping between the ID of the verse and its index in the list (thus verse with ID 5634 is located at index 5634 in the Numpy file). Verses not available for a given language (as stated by the value "Not Available" in the CSV file) are represented by empty lists in the Numpy files, thus ensuring a perfect verse-to-verse alignement between each file.

Spectrogram were extracted using Librosa with the following parameters:

Pre-emphasis = 0.97
Sample rate = 16000
Window size = 0.025
Window stride = 0.01
Window type = 'hamming'
Mel coefficients = 40
Min frequency = 20

Files

verses.csv

Files (29.5 GB)

Name	Size	Download all
basque_mel_spec.npy md5:3e9f5aa1baf62a3b9080c9ac120d332b	3.8 GB	Download
english_mel_spec.npy md5:a5e4e731fc37de7db243d852beafc977	3.2 GB	Download
finnish_mel_spec.npy md5:3209d995c7d5ce6087a86e9ad2bf84b0	4.0 GB	Download
french_mel_spec.npy md5:2bf7c075822ddf1eb9123d1c54b3c566	3.4 GB	Download
hungarian_mel_spec.npy md5:145d05246f1c1e5148045a62acd74a29	3.6 GB	Download
romanian_mel_spec.npy md5:7325bc93c989444716c907cdf002af30	4.0 GB	Download
russian_mel_spec.npy md5:c794e25ef60791b8ac1e326aafd7c534	3.8 GB	Download
spanish_mel_spec.npy md5:9ac66e413c5124243bdf949be1ac77a1	3.7 GB	Download
verses.csv md5:fcc0e45122138ac92ed793785991d5b9	2.9 MB	Preview Download

	All versions	This version
Views	602	601
Downloads	335	334
Data volume	1.2 TB	1.2 TB

MaSS - Multilingual corpus of Sentence-aligned Spoken utterances

Creators

Description

Files

verses.csv

Files (29.5 GB)