Fine-tuning vs. Scratch Training of wav2vec 2.0 for Flemish Dutch Speech Recognition
Description
Recent research in speech processing exhibits a growing interest in unsupervised and self-supervised representation learning from unlabelled data to alleviate the need for large amounts of annotated data. We investigate several popular pre-training methods and apply them to Flemish Dutch. We compare off-the-shelf English pre-trained models to models trained on an increasing amount of Flemish data. We find that the most important factors for positive transfer to downstream speech recognition tasks include a substantial amount of data and a matching pre-training domain. Ideally, we also finetune
Research goal: How does the fine-tuning of a pre-trained English wav2vec 2.0 model on Flemish Dutch data of varying sizes compare to training a model from scratch on the Common Voice Flemish Dutch dataset in terms of Word Error Rate (WER) and model convergence speed?
Autonomous synthesis report generated by Assignee Research. Tribunal consensus score: 7.5/10.
Notes
Files
paper.pdf
Files
(84.4 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:47173ee39674a3f9f0d03a320ec35fa2
|
84.4 kB | Preview Download |