Published May 30, 2025
| Version v2
Dataset
Open
HiACC: Hinglish Adult & Children Code-switched corpus
Authors/Creators
Contributors
Supervisor (2):
Description
The HiACC corpus is a novel Hinglish code-switched speech dataset featuring both adult and child speakers. It captures naturalistic code-switching through spontaneous responses to everyday questions, story reading, and image-based prompts. The dataset comprises 5.24 hours of segmented audio, including 3,318 utterances from adults and 1,858 from children, all of which have been manually transcribed and annotated for code-switching.
Files
Corpus.zip
Additional details
Dates
- Submitted
-
2025-05-27The HiACC corpus is a richly annotated Hinglish code-switched speech dataset featuring both adult and child speakers, designed for researchers working on code-switching, speech recognition, speaker analysis, and related NLP/ASR tasks. Each category is structured identically to maintain uniformity and support streamlined data loading, training, and analysis workflows.