Published November 30, 2025 | Version v1

Style-Controlled VALL-E for Few-Shot Emotional German TTS

  • 1. ROR icon Budapest University of Technology and Economics

Description

We present an expressive German Text-to-Speech (TTS) system built on a modified VALL-E neural codec language model. Our approach introduces a style-conditioning mechanism to enable emotional and prosodic control during speech synthesis. Using the Thorsten German Emotional TTS dataset, we preprocess and augment 2,400 utterances across 8 emotions. Our model integrates emotion tokens and style embeddings to guide expressive generation without explicit supervision. Preliminary results suggest that this method can produce prosodically varied and natural-sounding German speech, demonstrating its potential in low-resource emotional TTS settings.

Files

Style-Controlled_VALL-E_for_Few-Shot_Emotional_German_TTS.pdf

Files (2.7 MB)