Published March 25, 2024 | Version v1
Dataset Open

DCASE 2024 Challenge Task 7 Development Dataset : Environmental Sound Scene Synthesis

  • 1. Gaudio Lab inc.
  • 2. Doshita University
  • 3. ROR icon Laboratoire des Sciences du Numérique de Nantes
  • 4. Gaudio Lab, inc.
  • 5. ROR icon New York University
  • 6. ROR icon Ritsumeikan University

Description

Description

This dataset comprises embeddings and captions utilized as the development dataset for DCASE 2024 Challenge Task 7, focusing on 'Environmental Sound Scene Synthesis.' The embeddings are derived from 60 different 4-second audio files formatted as mono 32-bit 32kHz, and are contained in the 'embeddings.tar.xz' file. Captions corresponding to each audio file can be found in 'caption.csv'. This dataset does not comprise the audio files, only the embeddings. Three different types of embeddings are provided: VGGish (vggish), MS-CLAP (clap-2023), and PANNs CNN14 Wavegram-Logmel (panns-wavegram-logmel). Only PANNs CNN14 Wavegram-Logmel (panns-wavegram-logmel) embeddings are used for evaluation in the challenge. For further details, please refer to the challenge website.

Contact

  • Modan Tailleur, modan.tailleur@ls2n.fr
  • Mathieu Lagrange, mathieu.lagrange@ls2n.fr

Files

captions.csv

Files (545.3 kB)

Name Size Download all
md5:175045379f858ab70f5542e4837fffab
3.9 kB Preview Download
md5:9b599d9e1e83e07c461fe01673ac48c7
541.4 kB Download