scieee AI-readable full text Open interactive document viewer

Poster for "Image Sonification as Unsupervised Domain Transfer"

Laczkó, Bálint; Rognes, Marie E.; Jensenius, Alexander Refsum

Full text

Image Sonification as Unsupervised Domain Transfer Bálint Laczkó1[0009—0001—8337—2509], Marie E. Rognes2,3[0000—0002—6872—3710], and Alexander Refsum Jensenius1[0000—0001—6171—8743] 1 RITMO Centre for Interdisciplinary Studies in Rhythm, Time and Motion, Department of Musicology, University of Oslo, Norway 2 Dept. Numerical Analysis and Scientific Computing, Simula Research Laboratory, Norway; 3 K. G. Jebsen Centre for Brain Fluids Research, Oslo, Norway MAE (per axis) Image MSE + KLD(z) + KLD(ẑ) Image MLP Image FactorVAE Sound Encoder Z Decoder MSE + KLD(z) + KLD(ẑ) Sound Sound FactorVAE MLP Mapper Cycle Consistency (MSE) μ σ Encoder Z Decoder μ σ The proposed model architecture. The blue group denotes the image model, the green group the sound model, and the orange block represents the Mapper. Red arrows indicate loss criteria. The semi-transparent red path indicates how images are mapped to sound at test time. Learning Disentangled Representations Top row: Samples from the image dataset (left) and the latent space of the encoded images colored by the white square's X position (middle) and Y position (right). Bottom row: Samples from the sound dataset (left) and the latent space of the encoded sounds colored by the MIDI Pitch (middle) and Amplitude (right) parameters. Project repository: github.com/balintlaczko/cmmr25_isudt RITMO Centre for Interdisciplinary Studies in Rhythm, Time and Motion UiO Life Science