Quantize & Factorize: A Fast Yet Effective Unsupervised Audio Representation Without Deep Learning

Jaehun Kim; Matthew C. McCallum; Andreas F. Ehmann

doi:10.5281/zenodo.17706308

There is a newer version of the record available.

Published September 21, 2025 | Version v1

Conference paper Open

Quantize & Factorize: A Fast Yet Effective Unsupervised Audio Representation Without Deep Learning

Foundation models have become increasingly prevalent in tackling Music Information Retrieval (MIR) tasks. Although they can be a powerful tool for understanding music, the computation required for the training and inference of these models continues to grow as they become more complex. Specialized acceleration, such as Graphical Processing Units (GPUs), has become necessary for operating these models, as they are mostly based on large Deep Learning (DL) architectures. Furthermore, it is difficult for users to interpret them due to their black-box nature. In this work, we propose Quantizers and Factorizers for Music embeddings (QFM), a fast, unsupervised audio representation for music understanding backed by a wide range of rich MIR features and efficient feature learners. Experimental results show that QFM models perform within the range of results achieved by recent previous open source DL models on all evaluated tasks, with competitive results on a subset. This is surprising given the significantly smaller computational requirements of QFM models for training and inference.

Files

000092.pdf

Files (299.4 kB)

Name	Size	Download all
000092.pdf md5:ec299ac095a9a6befcc2f6b39e850e2b	299.4 kB	Preview Download

165

Views

Downloads

Show more details

	All versions	This version
Views	165	104
Downloads	80	50
Data volume	25.1 MB	15.6 MB

More info on how stats are collected....

DOI

Resource type

Conference paper

Publisher

ISMIR

Imprint

Proceedings of the 26th International Society for Music Information Retrieval Conference, 801-810. Daejeon, South Korea.

Conference

International Society for Music Information Retrieval Conference (ISMIR 2025) , Daejeon, South Korea and Online, September 21-25, 2025

License: Creative Commons Attribution 4.0 International

The Creative Commons Attribution license allows re-distribution and re-use of a licensed work on the condition that the creator is appropriately credited. Read more

Technical metadata

Created: November 25, 2025
Modified: November 25, 2025

Quantize & Factorize: A Fast Yet Effective Unsupervised Audio Representation Without Deep Learning

Authors/Creators

Description

Files

000092.pdf

Files (299.4 kB)