{
  "DOI": "10.5281/zenodo.3612637",
  "abstract": "FSDKaggle2019 is an audio dataset containing 29,266 audio files annotated with 80 labels of the AudioSet Ontology. FSDKaggle2019 has been used for the DCASE Challenge 2019 Task 2,\u00a0 which was run as a Kaggle competition titled\u00a0Freesound Audio Tagging 2019.\n\n\nCitation\n\n\nIf you use the FSDKaggle2019 dataset or part of it, please cite our DCASE 2019 paper:\n\n\n\n\nEduardo Fonseca, Manoj Plakal, Frederic Font, Daniel P. W. Ellis, Xavier Serra. \"Audio tagging with noisy labels and minimal supervision\". Proceedings of the DCASE 2019 Workshop, NYC, US (2019)\n\n\n\nYou can also consider citing our ISMIR 2017 paper, which describes how we gathered the manual annotations included in FSDKaggle2019.\n\n\n\n\nEduardo Fonseca, Jordi Pons, Xavier Favory, Frederic Font, Dmitry Bogdanov, Andres Ferraro, Sergio Oramas, Alastair Porter, and Xavier Serra, \"Freesound Datasets: A Platform for the Creation of Open Audio Datasets\", In Proceedings of the 18th International Society for Music Information Retrieval Conference, Suzhou, China, 2017\n\n\n\nData curators\n\n\nEduardo Fonseca, Manoj Plakal, Xavier Favory, Jordi Pons\n\n\nContact\n\n\nYou are welcome to contact Eduardo Fonseca should you have any questions at eduardo.fonseca@upf.edu.\n\n\n\u00a0\n\n\nABOUT FSDKaggle2019\n\n\nFreesound Dataset Kaggle 2019 (or FSDKaggle2019 for short) is an audio dataset containing 29,266 audio files annotated with 80 labels of the AudioSet Ontology [1]. FSDKaggle2019 has been used for the Task 2 of the Detection and Classification of Acoustic Scenes and Events (DCASE) Challenge 2019. Please visit the DCASE2019 Challenge Task 2 website for more information. This Task was hosted on the Kaggle platform as a competition titled Freesound Audio Tagging 2019. It was organized by researchers from the Music Technology Group (MTG) of Universitat Pompeu Fabra (UPF), and from Sound Understanding team at Google AI Perception. The competition intended to provide insight towards the development of broadly-applicable sound event classifiers able to cope with label noise and minimal supervision conditions.\n\n\nFSDKaggle2019 employs audio clips from the following sources:\n\n\n\n\t\n\u00a0Freesound Dataset (FSD): a dataset being collected at the MTG-UPF based on Freesound content organized with the AudioSet Ontology\n\t\n\u00a0The soundtracks of a pool of Flickr videos taken from the Yahoo Flickr Creative Commons 100M dataset (YFCC)\n\n\n\nThe audio data is labeled using a vocabulary of 80 labels from Google\u2019s AudioSet Ontology [1], covering diverse topics: Guitar and other Musical Instruments, Percussion, Water, Digestive, Respiratory sounds, Human voice, Human locomotion, Hands, Human group actions, Insect, Domestic animals, Glass, Liquid, Motor vehicle (road), Mechanisms, Doors, and a variety of Domestic sounds. The full list of categories can be inspected in\u00a0vocabulary.csv (see Files & Download below). The goal of the task was to build a multi-label audio tagging system that can predict appropriate label(s) for each audio clip in a test set.\n\n\nWhat follows is a summary of some of the most relevant characteristics of FSDKaggle2019. Nevertheless, it is highly recommended to read our DCASE 2019 paper for a more in-depth description of the dataset and how it was built.\n\n\nGround Truth Labels\n\n\nThe ground truth labels are provided at the clip-level, and express the presence of a sound category in the audio clip, hence can be considered\u00a0weak\u00a0labels or tags. Audio clips have variable lengths (roughly from 0.3 to 30s).\n\n\nThe audio content from\u00a0FSD\u00a0has been manually labeled by humans following a data labeling process using the\u00a0Freesound Annotator\u00a0platform. Most labels have inter-annotator agreement but not all of them. More details about the data labeling process and the\u00a0Freesound Annotator\u00a0can be found in [2].\n\n\nThe\u00a0YFCC\u00a0soundtracks were labeled using automated heuristics applied to the audio content and metadata of the original Flickr clips. Hence, a substantial amount of label noise can be expected. The label noise can vary widely in amount and type depending on the category, including in- and out-of-vocabulary noises. More information about some of the types of label noise that can be encountered is available in [3].\n\n\nSpecifically, FSDKaggle2019 features\u00a0three types of label quality, one for each set in the dataset:\n\n\n\n\t\ncurated train set: correct (but potentially incomplete) labels\n\t\nnoisy train set: noisy labels\n\t\ntest set: correct and complete labels\n\n\n\nFurther details can be found below in the sections for each set.\n\n\nFormat\n\n\nAll audio clips are provided as uncompressed PCM 16 bit, 44.1 kHz, mono audio files.\n\n\n\u00a0\n\n\nDATA SPLIT\n\n\nFSDKaggle2019 consists of\u00a0two train sets\u00a0and\u00a0one test set. The idea is to limit the supervision provided for training (i.e., the manually-labeled, hence reliable, data), thus promoting approaches to deal with label noise.\n\n\nCurated train set\n\n\nThe\u00a0curated train set\u00a0consists of manually-labeled data from\u00a0FSD.\u00a0\n\n\n\n\t\nNumber of clips/class: 75 except in a few cases (where there are less)\n\t\nTotal number of clips: 4970\n\t\nAvg number of labels/clip: 1.2\n\t\nTotal duration: 10.5 hours\n\n\n\nThe duration of the audio clips ranges from 0.3 to 30s due to the diversity of the sound categories and the preferences of Freesound users when recording/uploading sounds. Labels are correct but potentially incomplete. It can happen that a few of these audio clips present additional acoustic material beyond the provided ground truth label(s).\n\n\nNoisy train set\n\n\nThe\u00a0noisy train set\u00a0is a larger set of noisy web audio data from Flickr videos taken from the\u00a0YFCC\u00a0dataset [5].\n\n\n\n\t\nNumber of clips/class: 300\n\t\nTotal number of clips: 19,815\n\t\nAvg number of labels/clip: 1.2\n\t\nTotal duration: ~80 hours\n\n\n\nThe duration of the audio clips ranges from 1s to 15s, with the vast majority lasting 15s. Labels are automatically generated and purposefully noisy. No human validation is involved. The label noise can vary widely in amount and type depending on the category, including in- and out-of-vocabulary noises.\n\n\nConsidering the numbers above, the per-class data distribution available for training is, for most of the classes, 300 clips from the noisy train set and 75 clips from the curated train set. This means 80% noisy / 20% curated at the clip level, while at the duration level the proportion is more extreme considering the variable-length clips.\n\n\nTest set\n\n\nThe\u00a0test set\u00a0is used for system evaluation and consists of manually-labeled data from\u00a0FSD.\u00a0\n\n\n\n\t\nNumber of clips/class: between 50 and 150\n\t\nTotal number of clips: 4481\n\t\nAvg number of labels/clip: 1.4\n\t\nTotal duration: 12.9 hours\n\n\n\nThe acoustic material present in the test set clips is labeled exhaustively using the aforementioned vocabulary of 80 classes. Most labels have inter-annotator agreement but not all of them. Except human error, the label(s) are correct and complete considering the target vocabulary; nonetheless, a few clips could still present additional (unlabeled) acoustic content out of the vocabulary.\n\n\nDuring the\u00a0DCASE2019 Challenge Task 2, the test set was split into two subsets, for the\u00a0public\u00a0and\u00a0private\u00a0leaderboards, and only the data corresponding to the\u00a0public\u00a0leaderboard was provided.\u00a0In this current package you will find the full test set with all the test labels. To allow comparison with previous work, the file\u00a0test_post_competition.csv\u00a0includes a flag to determine the corresponding leaderboard (public or private) for each test clip (see more info in\u00a0Files & Download\u00a0below).\n\n\nAcoustic mismatch\n\n\nAs mentioned before, FSDKaggle2019 uses audio clips from two sources:\n\n\n\n\t\nFSD: curated train set and test set, and\n\t\nYFCC: noisy train set.\n\n\n\nWhile the sources of audio (Freesound and Flickr) are collaboratively contributed and pretty diverse themselves, a certain acoustic mismatch can be expected between\u00a0FSD\u00a0and\u00a0YFCC. We conjecture this mismatch comes from a variety of reasons.\nFor example, through acoustic inspection of a small sample of both data sources, we find a higher percentage of high quality recordings in FSD. In addition, audio clips in Freesound are typically recorded with the purpose of capturing audio, which is not necessarily the case in YFCC.\n\n\nThis mismatch can have an impact in the evaluation, considering that most of the train data come from YFCC, while all test data are drawn from FSD. This constraint (i.e., noisy training data coming from a different web audio source than the test set) is sometimes a real-world condition.\n\n\n\u00a0\n\n\nLICENSE\n\n\nAll clips in FSDKaggle2019 are released under Creative Commons (CC) licenses. For attribution purposes and to facilitate attribution of these files to third parties, we include a mapping from the audio clips to their corresponding licenses.\n\n\n\n\t\n\n\t\nCurated train set and test set. All clips in Freesound are released under different modalities of Creative Commons (CC) licenses, and each audio clip has its own license as defined by the audio clip uploader in Freesound, some of them requiring attribution to their original authors and some forbidding further commercial reuse. The licenses are specified in the files\u00a0train_curated_post_competition.csv\u00a0and\u00a0test_post_competition.csv. These licenses can be CC0, CC-BY, CC-BY-NC and CC Sampling+.\n\t\n\t\n\n\t\nNoisy train set. Similarly, the licenses of the soundtracks from\u00a0Flickr\u00a0used in FSDKaggle2019 are specified in the file\u00a0train_noisy_post_competition.csv. These licenses can be CC-BY and CC BY-SA.\n\t\n\n\n\nIn addition, FSDKaggle2019 as a whole is the result of a curation process and it has an additional license. FSDKaggle2019 is released under\u00a0CC-BY. This license is specified in the\u00a0LICENSE-DATASET\u00a0file downloaded with the\u00a0FSDKaggle2019.doc zip file.\n\n\n\u00a0\n\n\nFILES & DOWNLOAD\n\n\nFSDKaggle2019 can be downloaded as a series of zip files with the following directory structure:\n\n\nroot\n\u2502  \n\u2514\u2500\u2500\u2500FSDKaggle2019.audio_train_curated/               Audio clips in the curated train set\n\u2502\n\u2514\u2500\u2500\u2500FSDKaggle2019.audio_train_noisy/                 Audio clips in the noisy train set\n\u2502   \n\u2514\u2500\u2500\u2500FSDKaggle2019.audio_test/                        Audio clips in the full test set\n\u2502   \n\u2514\u2500\u2500\u2500FSDKaggle2019.meta/                              Files for evaluation setup\n\u2502   \u2502            \n\u2502   \u2514\u2500\u2500\u2500 train_curated_post_competition.csv          Ground truth for the curated train set\n\u2502   \u2502            \n\u2502   \u2514\u2500\u2500\u2500 train_noisy_post_competition.csv            Ground truth for the noisy train set\n\u2502   \u2502            \n\u2502   \u2514\u2500\u2500\u2500 test_post_competition.csv                   Ground truth for the full test set\n\u2502   \u2502            \n\u2502   \u2514\u2500\u2500\u2500 vocabulary.csv                              List of sound classes in FSDKaggle2019    \n\u2502   \n\u2514\u2500\u2500\u2500FSDKaggle2019.doc/\n    \u2502            \n    \u2514\u2500\u2500\u2500README.md                                    The dataset description file that you are reading\n    \u2502            \n    \u2514\u2500\u2500\u2500LICENSE-DATASET                              License of the FSDKaggle2019 dataset as an entity   \n\n\n\nImportant Note:\u00a0the original\u00a0train_curated.csv\u00a0and\u00a0train_noisy.csv\u00a0files provided during the competition have been updated with more metadata (licenses, Freesound/Flickr ids, etc.) into\u00a0train_curated_post_competition.csv\u00a0and\u00a0train_noisy_post_competition.csv. Likewise, the original\u00a0test.csv\u00a0that was not public during the competition is now available with ground truth and metadata as\u00a0test_post_competition.csv.\n\n\nEach row (i.e. audio clip) of the\u00a0train_curated_post_competition.csv\u00a0or\u00a0train_noisy_post_competition.csv files contains the following information:\n\n\n\n\t\nfname: the file name, e.g.,\u00a00006ae4e.wav\n\t\nlabels: the audio classification label(s) (ground truth). Note that the number of labels per clip can be one, eg,\u00a0Bark\u00a0or more, eg,\u00a0\"Walk_and_footsteps,Slam\".\n\t\nfreesound_id\u00a0or\u00a0flickr_video_URL: the Freesound id or Flickr id for the audio clip\n\t\nlicense: the license for the audio clip\n\n\n\nEach row (i.e. audio clip) of the\u00a0test_post_competition.csv\u00a0file contains the following information:\n\n\n\n\t\nfname: the file name\n\t\nlabels: the audio classification label(s) (ground truth). Note that the number of labels per clip can be one, eg,\u00a0Bark\u00a0or more, eg,\u00a0\"Walk_and_footsteps,Slam\".\n\t\nusage: string that indicates to which Kaggle leaderboard the clip was associated during the competition:\u00a0Public\u00a0or\u00a0Private\n\t\nfreesound_id: the Freesound id for the audio clip\n\t\nlicense: the license for the audio clip\n\n\n\nDetected corrupted files in the curated train set\n\n\nThe following 5 audio files in the\u00a0curated train set\u00a0have a wrong label, due to a bug in the file renaming process:\u00a0f76181c4.wav,\u00a077b925c2.wav,\u00a06a1f682a.wav,\u00a0c7db12aa.wav,\u00a07752cc8a.wav.\n\n\nThe audio file\u00a01d44b0bd.wav\u00a0in the\u00a0curated train set\u00a0was found to be corrupted (contains no signal) due to an error in format conversion.\n\n\nIf you find more corrupted files in FSDKaggle2019, please send an email to eduardo.fonseca@upf.edu.\n\n\nDownload\n\n\nEach of the folders in the directory structure above is compressed into one corresponding zip file that you can download and unzip with your favorite compression tool. There is one exception: due to the large size of\u00a0FSDKaggle2019.audio_train_noisy/, it is split into 7 files (note the last file is not\u00a0*.z07, but\u00a0*.zip):\n\n\nFSDKaggle2019.audio_train_noisy.z01\nFSDKaggle2019.audio_train_noisy.z02\nFSDKaggle2019.audio_train_noisy.z03\nFSDKaggle2019.audio_train_noisy.z04\nFSDKaggle2019.audio_train_noisy.z05\nFSDKaggle2019.audio_train_noisy.z06\nFSDKaggle2019.audio_train_noisy.zip\n\n\nIn this case, you first have to download the 7 files. Once downloaded, we convert the split archive to a single-file archive. In other words, we merge the 7 files into one zip file called e.g.\u00a0unsplit.zip\u00a0in your local machine.\n\n\nzip -s 0 FSDKaggle2019.audio_train_noisy.zip --out unsplit.zip\n\n\nFinally, this merged file is unzipped.\n\n\nunzip unsplit.zip\n\n\nBaseline System\n\n\nA CNN baseline system for FSDKaggle2019 is available at\u00a0https://github.com/DCASE-REPO/dcase2019task2baseline.\n\n\n\u00a0\n\n\nREFERENCES AND LINKS\n\n\n[1] Jort F Gemmeke, Daniel PW Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R Channing Moore, Manoj Plakal, and Marvin Ritter. \"Audio set: An ontology and human-labeled dataset for audio events.\" In Proceedings of the International Conference on Acoustics, Speech and Signal Processing, 2017. [PDF]\n\n\n[2] Eduardo Fonseca, Jordi Pons, Xavier Favory, Frederic Font, Dmitry Bogdanov, Andres Ferraro, Sergio Oramas, Alastair Porter, and Xavier Serra. \"Freesound Datasets: A Platform for the Creation of Open Audio Datasets.\" In Proceedings of the International Conference on Music Information Retrieval, 2017. [PDF]\n\n\n[3] Eduardo Fonseca, Manoj Plakal, Daniel P. W. Ellis, Frederic Font, Xavier Favory, and Xavier Serra. \"Learning Sound Event Classifiers from Web Audio with Noisy Labels.\" In Proceedings of the International Conference on Acoustics, Speech and Signal Processing, 2019. [PDF]\n\n\n[4] Frederic Font, Gerard Roma, and Xavier Serra. \"Freesound technical demo.\" Proceedings of the 21st ACM international conference on Multimedia, 2013.\u00a0https://freesound.org\n\n\n[5] Bart Thomee, David A. Shamma, Gerald Friedland, Benjamin Elizalde, Karl Ni, Douglas Poland, Damian Borth, and Li-Jia Li, YFCC100M: The New Data in Multimedia Research, Commun. ACM, 59(2):64\u201373, January 2016\n\n\nFreesound Annotator:\u00a0https://annotator.freesound.org/\nFreesound:\u00a0https://freesound.org\nEduardo Fonseca's personal website:\u00a0http://www.eduardofonseca.net/\nMore datasets collected by us:\u00a0http://www.eduardofonseca.net/datasets/\n\n\nAcknowledgments\n\n\nThis work is partially supported by the European Union\u2019s Horizon 2020 research and innovation programme under grant agreement No 688382\u00a0AudioCommons. Eduardo Fonseca is also sponsored by a\u00a0Google Faculty Research Award 2018. We thank everyone who contributed to FSDKaggle2019 with annotations.",
  "author": [
    {
      "family": "Eduardo Fonseca"
    },
    {
      "family": "Manoj Plakal"
    },
    {
      "family": "Frederic Font"
    },
    {
      "family": "Daniel P. W. Ellis"
    },
    {
      "family": "Xavier Serra"
    }
  ],
  "id": "3612637",
  "issued": {
    "date-parts": [
      [
        "2020",
        "01",
        "20"
      ]
    ]
  },
  "publisher": "Zenodo",
  "title": "FSDKaggle2019",
  "type": "dataset",
  "version": "1.0"
}