Published February 27, 2026 | Version v1
Video/Audio Embargoed

HiTZ-Aholab emotional speech synthesis dataset in Basque

  • 1. HiTZ Center - Aholab, University of the Basque Country UPV/EHU

Description

General description

This resource is a high-quality Basque emotional speech corpus compiled by HiTZ Zentroa / AhoLab.

It consists of studio-quality audio recordings in WAV format and their corresponding orthographic text transcriptions.

The corpus contains read speech produced by professional native Basque speakers, recorded in a professional studio under controlled acoustic conditions. In contrast to the previously released neutral corpus, this dataset focuses specifically on expressive emotional speech, carefully designed to provide balanced coverage of multiple basic emotions while maintaining high audio quality and pronunciation clarity.

The recordings were produced by two speakers, Maider and Antton, and cover a range of sentence types and orthographic patterns, including declarative, interrogative, and exclamative sentences, as well as Basque-specific spelling phenomena (e.g., “tt”).

Each utterance is labeled according to one of four emotion categories:

  • POZ_: poza (joy)

  • HAS_: haserre (anger)

  • HAR_: harridura (surprise)

  • TRI_: tristura (sadness)

The emotion label is encoded directly in the filename prefix.

Corpus composition

The corpus includes:

  • 16,000 utterances per speaker

  • 4,000 utterances per emotion category

  • 32,000 utterances in total

Technical details

Property Value
Language Basque (Euskara, eu)
Speakers Maider, Antton
Speaking style Read emotional speech
Recording Professional studio
Sample rate 48,000 Hz
Channels 1 (mono)
Audio format WAV

All recordings were produced under controlled acoustic conditions to ensure consistency and suitability for speech technology research.

In total, the corpus comprises:

  • Maider: 16,000 utterances, approximately 21 h 22 min

  • Antton: 16,000 utterances, approximately 22 h 36 min

The material has been designed to ensure:

  • Balanced emotional coverage

  • Diversity of sentence modalities

  • Inclusion of specific orthographic and phonetic patterns relevant to Basque

Intended use

This corpus was designed to support research and development in areas such as emotional or expressive text-to-speech (TTS), particularly for high-quality or neural TTS models that benefit from:

  • Clean, studio-recorded audio

  • Consistent speaking style

  • Accurate orthographic transcriptions

Data organization

The corpus is distributed as one compressed TAR archive per speaker, each containing the corresponding audio recordings in WAV format.

Within each archive:

  • Audio files are named using an emotion prefix and a unique identifier, e.g.:

    • POZ_00001.wav

    • HAS_00001.wav

    • HAR_00001.wav

    • TRI_00001.wav

  • All recordings within an archive correspond to utterances produced by a single speaker.

In addition, a plain-text transcription file is provided, as both speakers read the same sentences. Each line associates an utterance identifier with its orthographic transcription using the following format:

POZ_00001 text of the sentence

The utterance identifier matches the WAV filename (without extension), enabling straightforward pairing of audio files and transcriptions.

Licensing

Creative Commons Attribution 4.0 International (CC BY 4.0)

Ethical considerations

All speakers provided informed consent for the recording and distribution of their voices.

Funding

The development of this resource has been funded by the Ministerio para la Transformación Digital y de la Función Pública and Plan de Recuperación, Transformación y Resiliencia - Funded by EU – NextGenerationEU within the framework of the project ILENIA with reference 2022/TL22/00215335, and by a grant from the Department of Culture and Language Policy of the Basque Government (IKER-GAITU project).

Versioning

This is version 1.0 of the dataset.

Contact

aholab@aholab.ehu.eus

HiTZ Center - Aholab, University of the Basque Country UPV/EHU

https://aholab.ehu.eus/aholab/

https://www.hitz.eus/

 

Files

Embargoed

The files will be made publicly available on March 2, 2027.

Reason: This dataset is under a 12-month embargo to allow the research team to complete ongoing analyses and publish the primary scientific results derived from these data. Full access will be granted automatically upon expiration of the embargo period.