Published September 21, 2025
| Version v1
Conference paper
Open
Singing Voice Separation From Carnatic Music Mixtures Using a Regression-Guided Latent Diffusion Model
Authors/Creators
Description
Score-based diffusion models have demonstrated promise to separate individual sources from music mixture signals in a generative fashion, paving the way for a new class of solutions for this challenging task. However, existing works rely on clean multi-stem data, which is scarce for several repertoires, consequently compromising generalization. In this work, we explore the potential of generative modeling to perform weakly-supervised singing voice separation for Carnatic Music, a music repertoire for which large quantities of multi-stem recordings with bleeding between sources have been directly collected from live performances. We pre-train a latent diffusion model to perform preliminary separation of Carnatic vocals conditioned on the corresponding mixture. Then, through a separately trained regressor - using a clean, smaller, and out-of-domain dataset - we estimate the level of bleeding in the preliminary separations and guide the diffusion model toward generating cleaner samples. Albeit introducing artifacts, operating on a latent space allows for an efficient development of the system using limited computational resources. The objective and perceptual evaluations show the potential of latent diffusion together with regression guidance for weekly-supervised separation.
Files
000097.pdf
Files
(598.9 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:3ea26b853d3179b3354a579bfba27f98
|
598.9 kB | Preview Download |