DCASE 2024 Challenge Task 7 Development Dataset : Environmental Sound Scene Synthesis

Choi, Keunwoo; Heller, Laurie M.; Imoto, Keisuke; Lagrange, Mathieu; Lee, Junwon; McFee, Brian; Okamoto, Yuki; Tailleur, Modan

doi:10.5281/zenodo.10869644

Published March 25, 2024 | Version v1

Dataset Open

DCASE 2024 Challenge Task 7 Development Dataset : Environmental Sound Scene Synthesis

1. Gaudio Lab inc.
2. Doshita University
3. Laboratoire des Sciences du Numérique de Nantes
4. Gaudio Lab, inc.
5. New York University
6. Ritsumeikan University

Description

This dataset comprises embeddings and captions utilized as the development dataset for DCASE 2024 Challenge Task 7, focusing on 'Environmental Sound Scene Synthesis.' The embeddings are derived from 60 different 4-second audio files formatted as mono 32-bit 32kHz, and are contained in the 'embeddings.tar.xz' file. Captions corresponding to each audio file can be found in 'caption.csv'. This dataset does not comprise the audio files, only the embeddings. Three different types of embeddings are provided: VGGish (vggish), MS-CLAP (clap-2023), and PANNs CNN14 Wavegram-Logmel (panns-wavegram-logmel). Only PANNs CNN14 Wavegram-Logmel (panns-wavegram-logmel) embeddings are used for evaluation in the challenge. For further details, please refer to the challenge website.

Contact

Modan Tailleur, modan.tailleur@ls2n.fr
Mathieu Lagrange, mathieu.lagrange@ls2n.fr

Files

captions.csv

Files (545.3 kB)

Name	Size	Download all
captions.csv md5:175045379f858ab70f5542e4837fffab	3.9 kB	Preview Download
embeddings.tar.xz md5:9b599d9e1e83e07c461fe01673ac48c7	541.4 kB	Download

	All versions	This version
Views	768	768
Downloads	803	803
Data volume	126.6 MB	126.6 MB

DCASE 2024 Challenge Task 7 Development Dataset : Environmental Sound Scene Synthesis

Authors/Creators

Description

Description

Files

captions.csv

Files (545.3 kB)