Joint Training of Speech Enhancement and Speaker Verification for Low-SNR Multimodal Benchmarks
Description
Recent advancements in speaker verification techniques show promise, but their performance often deteriorates significantly in challenging acoustic environments. Although speech enhancement methods can improve perceived audio quality, they may unintentionally distort speaker-specific information, which can affect verification accuracy. This problem has become more noticeable with the increasing use of generative deep neural networks (DNNs) for speech enhancement. While these networks can produce intelligible speech even in conditions of very low signal-to-noise ratio (SNR), they may also sever
Research goal: Can joint training of speech enhancement and speaker verification modules mitigate information distortion compared to cascaded systems in low-SNR multimodal benchmarks?
Autonomous synthesis report generated by SOVEREIGN Research Kernel. Tribunal consensus score: 9.2/10.
Notes
Files
paper.pdf
Files
(77.5 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:e1e197b47b84489b12abedd42ae9759d
|
77.5 kB | Preview Download |