Zero-shot Singing Voice Conversion
Description
The task of Singing Voice Conversion(SVC) is to transform the voice of one singer(source) to someone else voice(target) yet preserving the content. Voice Conversion(VC) is well studied for speech but a relatively less explored topic in the context of the singing voice. While VC involves both changes in timbre and expression, modeled via fundamental frequency(F0) and timing; in the case of singing, the timing and pitch are constrained by a musical score and thereby, common algorithms designed for speech only, cannot be used directly to transform singing voice due to the lack of constraint on the melody in the system. We adapt a recently proposed state of the art system for speech VC to the case of singing VC. The system takes as input audio samples of the source and target singers and separates the content to preserve(lyrics, F0, and rhythm) from the singer-dependent content to transform(timbre). Then, the conversion is realized by re-synthesizing lyrics and melody of the source singer with the target singer’s timbre which is encoded using neural network embeddings. Thus, our system performs timbre conversion only and preserves the melody and expression from the source track. But we also try to perform zero-shot voice con-version which is a special case of VC when converting happens between voices that were not present in the training data i.e, conversion can be done between any two voices. We focus as well on how various algorithms for generating speaker/singer embeddings a˙ects the quality of our system in the context of zero-shot SVC.