Published March 3, 2024 | Version 1.1.1

CantiSong: EEG responses to continuous speech, song, and humming

  • 1. ROR icon Laboratoire des Systèmes Perceptifs
  • 2. ROR icon École Normale Supérieure - PSL
  • 3. EDMO icon University of Maryland
  • 4. ROR icon Trinity College Dublin

Description

This dataset contains the data used in the study Neural signatures of musical and linguistic interactions during natural song listening by G. Cantisani, S. Shamma, and G. Di Liberto. It includes EEG recordings collected while participants listened to 18 songs, together with their spoken and hummed versions.

The speech and song versions of each lyric were recorded by the same singer, while the humming was created by extracting the pitch contour from the original song and resynthesizing it with a voice-like sound. The stimuli were recorded by 6 singers (3 female, 3 male), with 3 songs per singer. The selected material comprises 20 minutes of speech, 43 minutes of sung speech, and 43 minutes of humming.

The .zip folder includes:

  • WAV audio files for all experimental stimuli;

  • Lyrics annotations for the spoken and sung materials;

  • Phoneme-level annotations with temporal alignments (manually annotated);

  • MIDI annotations of the melody (automatically annotated and manually corrected)
  • EEG recordings collected while participants listened to the stimuli.

The WAV files contain two channels: the first channel contains the auditory stimulus, while the second channel contains the trigger signal used to synchronize stimulus presentation with the EEG recording.

The original speech and singing recordings, together with the phoneme-level annotations, are derived from the NHSS Speech and Singing Parallel Database [1]. We gratefully acknowledge the authors for making these materials available and for granting permission to redistribute them as part of this dataset.

MIDI musical scores were obtained automatically using a note segmentation algorithm specific to singing voice transcription that converts a pitch sequence into a set of discrete and quantized note events [2]. We then visually inspect and manually correct the annotations.

The EEG recordings are released in the CND format; please check this webpage for further details. The shared data is preprocessed as in the study Neural signatures of musical and linguistic interactions during natural song listening by G. Cantisani, S. Shamma, and G. Di Liberto. Specifically, we bandpass-filtered between 1-8 Hz and downsampled to 64 Hz. 

Please refer to the accompanying paper for detailed information on the experimental protocol, stimulus construction, EEG acquisition and preprocessing, and data analysis procedures.

References

[1] Sharma, S., et al. (2021). NHSS: A Speech and Singing Parallel Database. Speech communication

[2] Mauch, M., and Dixon, S. (2014). Pyin: A fundamental frequency estimator using probabilistic threshold distributions. In 2014 IEEE Int. Conf. on acoustics, speech and signal processing (ICASSP)

Files

CantiSong-CND.zip

Files (3.4 GB)

Name Size
md5:64caca2c281bf1da1dc7fb4c459fc05a
3.4 GB Preview Download

Additional details

Related works

Is described by
Preprint: https://hal.science/hal-04529950 (URL)
Conference paper: 10.21437/Interspeech.2023-1949 (DOI)

Funding

European Commission
NEUME - Neuroplasticity and the Musical Experience 787836
Agence Nationale de la Recherche
FrontCog - Frontières en cognition ANR-17-EURE-0017
Science Foundation Ireland
13/RC/2106\_P2
NIH
R01-DC005779