Unifying Continuous and Discrete Compressed Representations of Audio

Marco Pasini; Stefan Lattner; George Fazekas

doi:10.5281/zenodo.17706477

There is a newer version of the record available.

Published September 21, 2025 | Version v1

Conference paper Open

Unifying Continuous and Discrete Compressed Representations of Audio

Efficiently representing audio signals in a compressed latent space is critical for latent generative modelling. However, existing autoencoders often force a choice between continuous embeddings and discrete tokens. Furthermore, achieving high compression ratios while maintaining audio fidelity remains a challenge. This paper introduces a novel audio autoencoder that overcomes these limitations by both efficiently encoding global features via summary embeddings, and by producing both compressed continuous embeddings at ~11 Hz and discrete tokens at a rate of 2.38 kbps from the same trained model, offering unprecedented flexibility for different downstream generative tasks. This is achieved through Finite Scalar Quantization (FSQ) and a novel FSQ-dropout technique, and does not require additional loss terms beyond the single consistency loss used for end-to-end training. Our model supports both autoregressive decoding and a novel parallel decoding strategy, with the latter achieving superior audio quality and faster decoding. Our model outperforms existing continuous and discrete autoencoders at similar bitrates in terms of reconstruction audio quality. Our work enables a unified approach to audio compression, bridging the gap between continuous and discrete generative modelling paradigms.

Files

000050.pdf

Files (400.8 kB)

Name	Size	Download all
000050.pdf md5:53e6e922ba8a7367de2bf5c57efe3759	400.8 kB	Preview Download

108

Views

170

Downloads

Show more details

	All versions	This version
Views	108	60
Downloads	170	151
Data volume	70.5 MB	62.5 MB

More info on how stats are collected....

DOI

Resource type

Conference paper

Publisher

ISMIR

Imprint

Proceedings of the 26th International Society for Music Information Retrieval Conference, 447-455. Daejeon, South Korea.

Conference

International Society for Music Information Retrieval Conference (ISMIR 2025) , Daejeon, South Korea and Online, September 21-25, 2025

License: Creative Commons Attribution 4.0 International

The Creative Commons Attribution license allows re-distribution and re-use of a licensed work on the condition that the creator is appropriately credited. Read more

Technical metadata

Created: November 25, 2025
Modified: November 25, 2025

Unifying Continuous and Discrete Compressed Representations of Audio

Authors/Creators

Description

Files

000050.pdf

Files (400.8 kB)