Acoustic Diversity Impact on Self-Supervised Speech Model Accuracy
Description
Self-supervised learning (SSL) has transformed speech processing, yet its reliance on massive pre-training datasets remains a bottleneck. While robustness is often attributed to scale and diversity, the role of the data distribution is less understood. We systematically examine how curated subsets of pre-training data influence Automatic Speech Recognition (ASR) performance. Surprisingly, optimizing for acoustic, speaker, or linguistic diversity yields no clear improvements over random sampling. Instead, we find that prioritizing the longest utterances achieves superior ASR results while using
Research goal: What is the effect of varying the acoustic diversity in pre-training data on the accuracy of self-supervised speech models, as measured by WER on LibriSpeech and Common Voice benchmarks?
Autonomous synthesis report generated by Assignee Research. Tribunal consensus score: 8.5/10.
Notes
Files
paper.pdf
Files
(83.7 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:5ade86dcfb14955f912667c65c08a0fa
|
83.7 kB | Preview Download |