Published June 12, 2026 | Version v3

TIMBRE: Layer-Wise Cross-Lingual Speech Emotion Recognition Across 49 Layers and 26 Corpora

Authors/Creators

  • 1. Independent Researcher

Description

Code and data to reproduce the Interspeech 2026 paper TIMBRE (paper #579). Run reproduce.py to regenerate every reported number from the included data: the CRIS layer (layer 15, 31% depth; cross-corpus

F1=0.392 with a 24.6% collapse at layer 45), within/between-type and -family contrasts with corpus-clustered bootstrap CIs, the acoustic vs. wav2vec2 mean-pool comparison (grand F1 0.2521 vs. 0.3553; acoustic features prevail on 3 of 26 corpora), and the acoustic-feature ablation. v1.2 adds the acoustic 26x26 cross-corpus matrix, replaces the Exp 2 mean-pool matrix with the recomputed final-dataversion (the superseded early-extraction matrix is kept for provenance), extends the corpus license table to all 26 corpora, and includes the supplementary PDF.

Base model: the public wav2vec2-xls-r-1b checkpoint on Hugging Face (not redistributed).

Corpus audio not redistributed; sources and licenses in CORPUS_SOURCES.md

Files

timbre-interspeech2026-v1.2.zip

Files (350.4 kB)

Name Size Download all
md5:cf3f06f32362e4fd64df82497b7c1263
350.4 kB Preview Download

Additional details

Related works

Software

Programming language
Python