Published July 31, 2025 | Version v1

Open-Domain Zero-Shot Audio Tagging: Evaluation via Semantic Embeddings

Authors/Creators

Contributors

  • 1. Universitat Pompeu Fabra

Description

This thesis investigates open-domain zero-shot audio tagging on the BSD10k dataset, a curated heterogeneous subset of Freesound, using Contrastive Language–Audio
Pretraining (CLAP) audio embeddings. To reduce the impact of rare and noisy labels, we apply a document frequency (DF) weighting scheme, which leads to substantial
performance gains. We further introduce a semantic evaluation approach based on SBERT text embeddings, which captures semantically valid tags missed by exact string matching. This yields notable gains across systems, with the largest improvements in the baseline model and consistent improvements for both the DFweighted variant and Freesound’s supervised tag recommender used for comparison. Together, the tag weighting and semantic evaluation demonstrate performance improvements beyond standard metrics. While the results show clear advances, zeroshot tagging with CLAP remains limited by incomplete generalization to folksonomy labels and sparse annotation coverage. Nevertheless, this work highlights the potential of zero-shot approaches to enable consistent and standardized audio annotation directly from raw audio.

Files

Tolga_Yapici_Master_Thesis_2025.pdf

Files (3.7 MB)

Name Size Download all
md5:e526c8d51269c8c4fcf6a4b3e65b0ff7
3.7 MB Preview Download

Additional details

Dates

Accepted
2025-10-09