scieee AI-readable full text Open interactive document viewer

EVxRAVE: Incorporating Neural Synthesis in an Augmented String Instrument Platform

Lindgren, Brian

Abstract

Extended abstract and poster for demonstration presentation at NeurIPS 2025 (San Diego) on 07 December, 2025.

Full text

EVxRAVE: Incorporating Neural Synthesis in an Augmented String Instrument Platform Brian Lindgren Composition and Computer Technologies (CCT) Department of Music, College of Arts and Sciences University of Virginia How can neural audio models extend the sonic potential of an augmented string instrument in real time? The EV’s integration with RAVE (Realtime Audio Variational autoEncoder) reframes the instrument’s own sound as a metaresonator: a system that listens to itself through the lens of a neural model. Instead of processing sound as a fixed effect, the model becomes a parallel pathway back into the instrument’s timbral identity, accessed through a different causal route. The latent space acts as a landscape of possibilities—each dimension a potential resonance between gesture, synthesis, and the acoustic body. the EV The EV is an augmented viola that merges acoustic performance with real-time digital signal processing. A quad infrared pickup captures each string’s motion independently, feeding the acoustic signal into a Bela single-board computer that runs C++ code to transmit performance data to six instances of Pure Data (Pd) on a Mac laptop, where the main processing and synthesis occur. The Pd environment is organized into multiple instances: one dedicated to each string, one managing central command and data routing, and one providing the graphical interface for performance control. RAVE integration EV signal flow RAVE is a neural audio synthesis model developed by Antoine Caillon and Philippe Esling at IRCAM*. Designed for high-quality, real-time performance, it learns a compact latent representation of sound *Caillon, A., & Esling, P. (2021). RAVE: A variational autoencoder for fast and high-quality neural audio synthesis. IRCAM – Sorbonne Université. Building the model through a two-stage training process: first as an autoencoder capturing perceptually meaningful spectral features, then through adversarial fine-tuning that enhances the naturalness of the generated signal. Trained on recordings of the EV, RAVE enables realtime synthesis in which latent dimensions can be mapped to musical parameters such as pitch, amplitude, or spectral features. Performing the model <<listen to the EV on YouTube