Published May 30, 2025 | Version v2

HiACC: Hinglish Adult & Children Code-switched corpus

Authors/Creators

Contributors

  • 1. ROR icon University of Ulster
  • 2. ROR icon University of Petroleum and Energy Studies

Description

The HiACC corpus is a novel Hinglish code-switched speech dataset featuring both adult and child speakers. It captures naturalistic code-switching through spontaneous responses to everyday questions, story reading, and image-based prompts. The dataset comprises 5.24 hours of segmented audio, including 3,318 utterances from adults and 1,858 from children, all of which have been manually transcribed and annotated for code-switching.

Files

Corpus.zip

Files (531.6 MB)

Name Size
md5:dd6cc9354e1dee5e2f25bc5243df88ac
531.6 MB Preview Download

Additional details

Dates

Submitted
2025-05-27
The HiACC corpus is a richly annotated Hinglish code-switched speech dataset featuring both adult and child speakers, designed for researchers working on code-switching, speech recognition, speaker analysis, and related NLP/ASR tasks. Each category is structured identically to maintain uniformity and support streamlined data loading, training, and analysis workflows.