IFA Dialog Video corpus
Contributors
Annotator (3):
Rights holder:
Supervisor (2):
Description
IFA Dialog Video corpus
All materials are licensed under the GNU general public license
© 2007, Nederlandse Taalunie
This corpus was made possible by grant 276-75-002 of
the Netherlands Organization for Scientific Research
Introduction
The IFA Dialog Video corpus is a collection of annotated video recordings of friendly Face-to-Face dialogs. It is modelled on the Face-to-Face dialogs in the Spoken Dutch Corpus (CGN). The procedures and design of the corpus were adapted to make this corpus useful for other researchers of Dutch speech. For this corpus, 24 dialog conversations of 15 minutes were recorded, in total 6 hours of speech. To stay close to the very useful Face-to-Face dialogs in the CGN, pairs of well acquainted participants were selected, either good friends, relatives, or long-time colleagues. The participants were allowed to talk about any topic they wanted.
In total, 20 of these conversations, 5 hours, were annotated to the same, or updated, standards as the original CGN. Only the initial orthographic transcription was done by hand. Other CGN-format annotations were only done automatically. Two other manual annotations were added, a functional annotation of dialog utterances and annotated gaze direction.
Recordings are coded by camera (DVA or DVB), converstation number (1-24, recordings 5, 18, 21, and 23 were not annotated), and speaker ID (A-AP). For example, DVA16AA and DVB16AB are the two camera recordings made from conversation 16, with subject AA recorded on camera A and subject AB recorded on camera B. Some subjects participated in more than one conversation, always with different partners.
See also the LREC paper The IFADV corpus: A free dialog video corpus
(van Son, R., Wesseling, W., Sanders, E., and van den Heuvel, H. (2008). LREC'08, Marrakech)
Other annotations can be found in El Haddad (2025).
Recordings
Recordings were made with two gen-locked JVC TK-C1480B analog color video cameras between April 25 and August 9, 2006.
Specification:
- Image pickup: 1/2 type IT CCD 752 (H) x 582 (V)
- Synchronization: Internal Line Lock, Full Genlock
- Scanning frequency: (H) 15.625kHz x (V) 50Hz
- Resolution: 480 TV lines (H)
- A: Ernitec GA4V10NA-1/2 lens (4-10mm)
- B: Panasonic WV-LZ80/2 lens (6-12mm)
Gen-lock ensures synchronization of all frames of the two cameras. Recordings were digitized using two Canopus ADVC110 digital video converters. Recordings were stored unprocessed on disk, ie, in DV format with 48 kHz 16 bit PCM sound.
Each camera was positioned to the left of one speaker and focussed on the face of the other. Subjects wore a Samson QV head-set microphone.
Subjects first spoke some scripted sentences. Then they were instructed to speak freely while preferably avoiding sensitive material or identifying people by name. All subject signed an informed consent and transfered all copyrights to the Dutch Language Union (Nederlandse Taalunie).
Point your IMDI browser to the file 'Annotations/IMDI/IFADVcorpus.imdi' or:
http://www.fon.hum.uva.nl/IFA-SpokenLanguageCorpora/IFADVcorpus/Annotations/IMDI/IFADVcorpus.imdi
Materials
Release note:
The original recordings contained dropped frames which made the two recordings of each dialog become out-of-sync. This has been corrected by duplicating frames. This procedure is described in the SMILoverlay files. Only the corrected recordings are made available here. The original recordings, with the lacking frames, are available on request. Recordings are limited to 900 seconds (15 min) and corrected for dropped frames. That is, the video frames and audio files of both recordings are synchronized.
- Compressed video, with automatically normalized brightness and contrast levels
- AVI: MPEG-4 Part 2/DivX3, encoded video recordings (~287 MB). Also normalized for sound volume.
mencoder -quiet -af volnorm=2:0.25 -vf pp=autolevels:fullyrange -of avi -ovc lavc -lavcopts vcodec=msmpeg4:vbitrate=2400000:vhq:keyint=50 -oac mp3lame -o outfile.avi infile.dv;
- OGV: Ogg Theora encoded video recordings (~200 MB)
ffmpeg2theora --format dv --videoquality 4 --sharpness 1 --pp autolevels:fullyrange --license GPLv2 -o outfile.ogv infile.dv
- AVI: MPEG-4 Part 2/DivX3, encoded video recordings (~287 MB). Also normalized for sound volume.
- Speech files, extracted from the recordings.
Extensions- WAV: RIFF/WAV (uncompressed)
- FLAC: Free Lossless Audio Codec (compressed)
- Annotations of the stereo sound of the A recording (ie, subject A on the left channel B on the right channel)
Labeling was done by Anita van Boxtel at SPEX under supervision of Eric Sanders and Henk van den Heuvel
(see: van Son et al, 2008 for annotation labels)
- ort: Orthographic transcription (Praat TextGrid)
- awd: Automatic word and phoneme alignment of ort (Praat TextGrid)
- pos: Automatic Part-of-Speech labeling (CGN format)
- EAF: Elan annotations for gaze direction
g: gazing at the partner, x: looking away, d: looking down, k: blinking - dbl: Normalized combination of other annotations (Praat TextGrid)
- DVA*.ort.awd.dbl: Functional dialogue annotations
- DVA*.gaze.dbl: Gaze annotation
- DVA*.ainton.dbl: Automatically determined end intonation (1: low, 2: mid, 3: high)
- scripts: Maintenance scripts for annotations
- Transcripts: Readable dialog transcripts
Corresponding audio files can be found in the Speech files directory. - Summaries of the dialogs
Compiled by Stephanie Wagenaar - IMDI files of the recordings
Compiled by Maaike van Naerssen
- scripts
Scripts used to record and process the dialogs - SMILoverlays
SMIL xml files to correct the dropped frames in the recordings in DialogCorpus - tables
Metadata on the recordings and the speakers as well as tab-separated-values database tables with a single record for every annotated item in the Annotations folder - Documents
Forms and published papers. - Annotation instructions (HTML)
Annotation instructions (Praat Manual)
Annotation instructions are in Dutch
IFA Dialog Video Corpus.
Copyright © 2007 Nederlandse Taal Unie
This corpus was made possible by grant 276-75-002 of the Netherlands Organization for Scientific Research
'Integration of information in spoken communication'
Created by R.J.J.H. van Son and Wieneke Wesseling of the ACLC. Annotations were performed by SPEX
Please note that these materials are distributed under the the GNU Affero General Public License v3.0 or later license. This license only covers the Copyright protection of the corpus. These materials are shared under the expectation that they are used ethically and responsibly. Publishing or broadcasting of materials from this corpus might be covered by other laws, eg, laws protecting the privacy and "good name" of the subjects. This is especially relevant if the materials are used outside of an educational or R&D context. Please read the forms in the Documents directory for more information (in Dutch).
Methods
Participants
The corpus consists of 20 annotated dialogs (selected from 24 recordings). All participants signed an informed consent and transferred all copyrights to the Dutch Language Union (Nederlandse Taalunie). For two minors, the parents too signed the forms. In total 34 speakers participated in the annotated recordings: 10 male and 24 female. Age ranged from 21 to 72 for males and 12 to 62 for females. All were native speakers of Dutch.
Participants originated in different parts of the Netherlands. Each speaker completed a form with personal characteristics. Notably, age, place of birth, and the places of primary and secondary education were all recorded. In addition, the education of the parents and data on height and weight, were recorded, as well as some data on training or experiences in relevant speech related fields, like speech therapy, acting, and call-center work.
The recordings were made in-face with only a small off-set (see Figure 3 in van Son et al. 2008). Video recordings were synchronized to make uniform timing measurements possible. All conversations were ”informal” since participants were friends or colleagues. There were no constraints on subject matter, style, or other aspects. However, participants were reminded before the recordings started that their speech would be published.
Table of contents
Name MD5 Size
# Annotation files and scripts
Annotations.zip md5:2d960d28b266e13e64f5db1d96cb4fe4 12.5 MB
# Audio files (FLAC)
AudioFLAC.zip md5:bf8d757eb21b114177a4af2d1b30d7d7 3.2 GB
# Audio files (WAV)
AudioWAV.zip md5:098f78ad5bd5f07b641df8ba53474d2a 5.8 GB
# Video cropped for side-by-side use in ELAN (AVI, MP3 audio)
CroppedMP3.zip md5:35e6e73ae0c8dd7ac12a595567ac3d04 14.1 GB
# General documents, forms, and articles
Documents.zip md5:79324df8b1f54e6e28243e09bbad4e23 1.5 MB
# Scripts used to process the video and audio recordings
scripts.zip md5:47f4c6914d207098fab92ceef29d5303 15.8 kB
# SMIL overlay files used to synchronize video recordings
SMILoverlays.zip md5:b010b68bf38f0f1ade3e49006d71b2c0 17.8 kB
# Tab-separated-values database tables (.txt) for all annotated items
tables.zip md5:27977e186d058d3d379e2d5ef1fbb67e 15.6 MB
# Video annotation instructions and manuals
VideoAnnotatie.zip md5:7cac54cd33b6067e1ec3dab73404229f 23.7 kB
# Video files (OGV, OGG Theora Video File), the recordings of conversation 21 are missing
VideoOGV.zip md5:2a7a251a876a1dcaa48c25af884edecd 8.7 GB
Files
Documents.zip
Files
(31.7 GB)
| Name | Size | |
|---|---|---|
|
md5:2d960d28b266e13e64f5db1d96cb4fe4
|
12.5 MB | Preview Download |
|
md5:bf8d757eb21b114177a4af2d1b30d7d7
|
3.2 GB | Preview Download |
|
md5:098f78ad5bd5f07b641df8ba53474d2a
|
5.8 GB | Preview Download |
|
md5:35e6e73ae0c8dd7ac12a595567ac3d04
|
14.1 GB | Preview Download |
|
md5:79324df8b1f54e6e28243e09bbad4e23
|
1.5 MB | Preview Download |
|
md5:47f4c6914d207098fab92ceef29d5303
|
15.8 kB | Preview Download |
|
md5:b010b68bf38f0f1ade3e49006d71b2c0
|
17.8 kB | Preview Download |
|
md5:27977e186d058d3d379e2d5ef1fbb67e
|
15.6 MB | Preview Download |
|
md5:7cac54cd33b6067e1ec3dab73404229f
|
23.7 kB | Preview Download |
|
md5:2a7a251a876a1dcaa48c25af884edecd
|
8.7 GB | Preview Download |
Additional details
Related works
- Is cited by
- Dataset: 10.5281/zenodo.15548946 (DOI)
- Conference proceeding: 10.21437/Interspeech.2025-1526 (DOI)
- Preprint: arXiv:2506.00981 (arXiv)
- Is described by
- Dataset: 10.5281/zenodo.14759144 (DOI)
Funding
- Dutch Research Council
- Integration of information in spoken communication 276-75-002
References
- van Son, R., Wesseling, W., Sanders, E., & van den Heuvel, H. (2008, May). The IFADV Corpus: a Free Dialog Video Corpus. In Proceedings of the Sixth International Conference on Language Resources and Evaluation (LREC'08).
- van Son, R.J.J.H., Wesseling, W., Sanders, E., van den Heuvel, H. (2009). Promoting free Dialog Video Corpora: The IFADV Corpus Example. In: Kipp, M., Martin, JC., Paggio, P., Heylen, D. (eds) Multimodal Corpora. MMCorp 2008. Lecture Notes in Computer Science(), vol 5509. Springer, Berlin, Heidelberg. https://doi.org/10.1007/978-3-642-04793-0_2
- El Haddad, K. (2025). Interaction Behavior Dataset [Data set]. Zenodo. https://doi.org/10.5281/zenodo.14759144