Published November 30, 2025
| Version v1
Style-Controlled VALL-E for Few-Shot Emotional German TTS
Authors/Creators
Description
We present an expressive German Text-to-Speech (TTS) system built on a modified VALL-E neural codec language model. Our approach introduces a style-conditioning mechanism to enable emotional and prosodic control during speech synthesis. Using the Thorsten German Emotional TTS dataset, we preprocess and augment 2,400 utterances across 8 emotions. Our model integrates emotion tokens and style embeddings to guide expressive generation without explicit supervision. Preliminary results suggest that this method can produce prosodically varied and natural-sounding German speech, demonstrating its potential in low-resource emotional TTS settings.
Files
Style-Controlled_VALL-E_for_Few-Shot_Emotional_German_TTS.pdf
Files
(2.7 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:c0a168687559627d0d5ec2be073fb7f5
|
2.7 MB | Preview Download |