ESC50Mix dataset
Description
Reference:
Jakob Abeßer, Sascha Grollmisch, Meinard Müller, How Robust are Audio Embeddings for Polyphonic Sound Event Tagging? (to be published in IEEE/ACM Transactions on Audio, Speech, and Language Processing)
The dataset builds upon the ESC-50 dataset (https://github.com/karolpiczak/ESC-50) and creates mixtures of two random sound pairs by systematically blending from one to the other sounds.
In total, the dataset includes 30,000 5s audio clips
- 1000 sound pairs
- 6 gain factors (used for mixing pairs)
- 5 data augmentations (none, loudness reduction, Gaussian noise, high-frequency boost, low-frequency boost)