There is a newer version of the record available.

Published September 21, 2025 | Version v1

Singing Voice Separation From Carnatic Music Mixtures Using a Regression-Guided Latent Diffusion Model

Description

Score-based diffusion models have demonstrated promise to separate individual sources from music mixture signals in a generative fashion, paving the way for a new class of solutions for this challenging task. However, existing works rely on clean multi-stem data, which is scarce for several repertoires, consequently compromising generalization. In this work, we explore the potential of generative modeling to perform weakly-supervised singing voice separation for Carnatic Music, a music repertoire for which large quantities of multi-stem recordings with bleeding between sources have been directly collected from live performances. We pre-train a latent diffusion model to perform preliminary separation of Carnatic vocals conditioned on the corresponding mixture. Then, through a separately trained regressor - using a clean, smaller, and out-of-domain dataset - we estimate the level of bleeding in the preliminary separations and guide the diffusion model toward generating cleaner samples. Albeit introducing artifacts, operating on a latent space allows for an efficient development of the system using limited computational resources. The objective and perceptual evaluations show the potential of latent diffusion together with regression guidance for weekly-supervised separation.

Files

000097.pdf

Files (598.9 kB)

Name Size Download all
md5:3ea26b853d3179b3354a579bfba27f98
598.9 kB Preview Download