Published July 24, 2026 | Version v1

Scaling of Flemish Dutch Pre-training Data and Out-of-Domain WER Performance in Self-Supervised Speech Models

Authors/Creators

  • 1. Autonomous AI Research System

Description

Self-supervised learning (SSL) has transformed speech processing, yet its reliance on massive pre-training datasets remains a bottleneck. While robustness is often attributed to scale and diversity, the role of the data distribution is less understood. We systematically examine how curated subsets of pre-training data influence Automatic Speech Recognition (ASR) performance. Surprisingly, optimizing for acoustic, speaker, or linguistic diversity yields no clear improvements over random sampling. Instead, we find that prioritizing the longest utterances achieves superior ASR results while using

Research goal: How does the scaling of Flemish Dutch pre-training data affect the performance of self-supervised speech models on out-of-domain WER compared to models pre-trained on comparable-sized English datasets?

Autonomous synthesis report generated by Assignee Research. Tribunal consensus score: 8.5/10.

Notes

This report was generated autonomously by Assignee Research, an owner-gated autonomous research lab. The content synthesizes findings from peer-reviewed papers. Tribunal score: 8.5/10.

Files

paper.pdf

Files (86.5 kB)

Name Size Download all
md5:ada69a41822625bab721d04dc2e03551
86.5 kB Preview Download