Published November 3, 2025 | Version v1

Image Sonification as Unsupervised Domain Transfer

Description

The process of image sonification maps visual features into perceived auditory features. Most established sonification methods rely on identifying salient visual features in the input data and then mapping their distribution to a proportional distribution of auditory features. However, this approach requires both domain expertise and manual feature engineering. Here, we propose a new method of image sonification, leveraging recent advances in representation learning and domain transfer. Our approach introduces a pair of variational auto-encoder models that learn disentangled latent representations of the images and sounds, respectively, and a separate network that maps between these representations. The resulting sonification system encodes images into the latent space and then decodes them as sounds. Both representations and their mapping are learned in an entirely unsupervised manner. When evaluating the system in an interactive real-time setting, we observed that the model successfully learned disentangled representations of image and sound factors in our synthetic datasets.

Files

CMMR2025_P1_15.pdf

Files (8.4 MB)

Name Size Download all
md5:575f7653463e4c6edddcd7306e505b3b
8.4 MB Preview Download