Published August 1, 2026 | Version 1.0.0

Zero-Latency Client-Side Audio Vocal Separation via Stereo Phase Subtraction and Formant Notch Filter Banks

Description

Deep-learning-based vocal separation architectures (e.g., Spleeter, Demucs) achieve impressive stem isolation but incur significant computational overhead, latency (10–30 seconds per track), and server infrastructure costs due to cloud processing. In this paper, we introduce a deterministic, zero-latency (0ms processing delay) client-side digital signal processing (DSP) framework for real-time vocal suppression running natively in modern web browsers via Web Audio API. Our algorithm integrates stereo phase subtraction (L - R), a 10-band surgical formant notch filter bank (300Hz - 3500Hz), parallel sub-bass (<220Hz) and high-frequency (>8kHz) preservation channels, and a 20ms Haas Effect psychoacoustic delay to reconstruct 3D stereo spatial width. Empirical benchmarks confirm 100% client-side execution, zero data transfer latency, and complete user data privacy.

Files

RESEARCH_PAPER.md

Files (43.0 kB)

Name Size Download all
md5:45308856d0d887663619a7290b7373b5
37.7 kB Download
md5:1f42f872d9877af6357f1ba678bacaa2
5.3 kB Preview Download

Additional details

Software