Published August 21, 2026 | Version v1

MaMa Sounds: A behaviourally normed database of 1,377 natural sounds for auditory cognition and neuroscience

  • 1. Institut des Neurosciences de La Timone, UMR 7289, CNRS and Aix-Marseille University, Marseille, France
  • 2. Department of Cognitive Neuroscience, Faculty of Psychology and Neuroscience, Maastricht University, Maastricht, Netherlands
  • 3. BISS institute, Faculty of Science and Engineering, Maastricht University, Maastricht, The Netherlands

Description

MaMa Sounds is a behaviourally normed database of 1,377 standardized natural sounds representing 240 expert-defined source-action classes. The resource combines standardized audio with detailed behavioural characterization from human listeners and is designed to support stimulus selection, matching, and continuous modelling in auditory cognition and neuroscience.

The sounds were curated from FSD50K and manually segmented from identified event onsets, trimmed or zero-padded to 2 s, level-normalized, and labelled with a noun identifying the sound-generating source and a verb identifying the action or mechanism.

About MaMa Sounds

The main characteristics of the dataset are:

  • 1,377 natural sounds spanning 240 expert-defined source-action classes;

  • standardized 2-second, mono, 16-bit PCM audio sampled at 16 kHz;

  • expert-defined noun and verb labels separately describing the sound-generating source and action or mechanism;

  • deidentified trial-level identification data from 365 participants, comprising 55,480 complete identification trials;

  • deidentified familiarity data from 275 participants, comprising 108,212 valid familiarity-rating trials;

  • sound-level behavioural norms for identification response time, number of playbacks, confidence, Word2Vec reference similarity, reference-retrieval percentile, between-listener agreement, and familiarity;

  • noun, verb, and joint noun-verb semantic norms;

  • direct mean and median estimates of the behavioural norms, together with the number of contributing observations;

  • two principal-component scores providing compact overall behavioural-identifiability measures;

  • participant and reference Word2Vec representations;

  • PCA parameters and validation outputs;

  • Python code for reproducing deterministic response cleaning, Word2Vec representations, sound-level norms, PCA scores, and validation outputs from the deposited deidentified participant data.

The individual behavioural measures are retained to support process-specific analyses, while the two overall behavioural-identifiability scores provide compact summaries of shared variation across response ease, semantic correspondence, agreement, and familiarity.

License

MaMa Sounds is released under the Creative Commons Attribution 4.0 International (CC BY 4.0) license.

The audio files included in MaMa Sounds are derived from clips distributed as part of FSD50K and originally uploaded to Freesound. Individual audio files retain the Creative Commons license associated with their corresponding source clip.

The Google News Word2Vec model weights are not redistributed.

Files

The complete public release is distributed as a ZIP archive. After extraction, its main contents are:

  • data/sounds/ contains the 1,377 standardized WAV files.

  • data/identification/ contains the 365 deidentified participant-level identification files, including trial index, sound filename, number of playbacks, original noun and verb responses, confidence, and identification response time.

  • data/identification/01_automated_cleanup/ contains the identification data with deterministic cleaned noun and verb fields and validity flags.

  • data/identification/02_w2v/ contains Word2Vec representations of cleaned participant responses, exact vocabulary keys, availability flags, and vocabulary-coverage information.

  • data/familiarity_ratings/ contains the 275 deidentified participant-level familiarity files.

  • data/MaMa_Sounds_reference_w2v_embeddings.tsv contains Word2Vec representations of the expert noun and verb labels.

  • data/MaMa_Sounds_table.tsv is the canonical 1,377-row master table, containing stimulus metadata and behavioural norms.

  • data/MaMa_Sounds_overall_norm_pca_parameters.tsv contains the parameters used to derive the overall behavioural-identifiability scores.

  • validation tables and correlation matrices describe associations among the behavioural norms.

  • code/ contains the public Python processing code used to reproduce deterministic response cleaning, participant and reference Word2Vec representations, sound-level behavioural norms, PCA scores, and validation outputs.

Please refer to the included README for detailed documentation of the archive contents, directory structure, file formats, variables, dependencies, and processing workflow.

Data protection

Public participant identifiers are experiment-specific and are not cross-linked between the identification and familiarity experiments.

Original platform exports, direct participant identifiers, demographics, private quality-control tables, excluded submissions, and private tests are not included in the public release.

Funding

This work was funded by the French National Research Agency (ANR-21-CE37-0027-01; ANR-16-CONV-0002 ILCB), the Dutch Research Council (NWO 406.20.GO.030), and the European Research Council through ERC-2024-SyG NASCE, Grant Agreement No. 101167313.

Contact

Bruno L. Giordano
Institut des Neurosciences de La Timone, UMR 7289, CNRS and Aix-Marseille University, Marseille, France
bruno.giordano@univ-amu.fr

Files

MaMa_Sounds_public.zip

Files (207.5 MB)

Name Size Download all
md5:529fc659445ae6974af68eac53063689
207.5 MB Preview Download

Additional details

Funding

European Commission
NASCE — Natural Auditory SCEnes in Humans and Machines: Establishing the Neural Computations of Everyday Hearing 101167313