Unified Timbre Transfer [paper] [code]

A Compact Model for Real-Time Multi-Instrument Sound Morphing

Anders R. Bargum, Naotake Masuda, Bogdan Teleaga, Andrew Fyfe and Cumhur Erkut

Abstract

Recent advances in transformer-and diffusion-based deep-generative models have significantly impacted the field of music and audio synthesis. However, controllable and real-time interactive models, such as those used for timbre transfer in music production, remain largely dominated by auto-encoders and generative adversarial networks. In pursuit of efficient and flexible timbre morphing and multi-instrument timbre transfer, we propose a simplified modeling approach, utilizing an upsampled two-dimensional timbre space in conjunction with engineered and instrument-dependent, excitation signals. Our model enables any-to-many timbre transfer with precise control over timbre, pitch, and loudness, while also allowing for seamless interpolation between instruments, eliminating the need for separate model training. We achieve performance comparable to specialized models, while supporting 44.1 kHz generation, making it highly relevant for the broader creative music community.


Pipeline Diagram

Background

We introduce a novel, monophonic, multi-instrument timbre-transfer framework that improves flexibility, control, and real-time performance while simplifying the generation process. Inspired by both DDSP and RAVE, our approach uses a generator conditioned on a perceptual timbre space, acting as a neural filter that processes an instrument-dependent excitation signal.

Live Demonstration

Left: We map 3 instruments (Clarinet, Trumpet, Violin) to discrete areas of a knob in the Neutone FX VST, the input is audio from a synthesiser. Right: We morph between the different instruments in real-time using audio samples as input (GUI made by Ze Hu).

Audio Comparison

We compare timbre transfer performance against two baseline approaches; VAE-GAN and P-RAVE. The timbre transfer is conducted with a violin and trumpet target. Inputs consist of URMP test samples and a subset of violin and trumpet inputs from the CocoChorales dataset.

Model Comparison
Target:
Input
Proposed
P-RAVE
VAE-GAN

Comparison of Model Architectures

The ablation study below demonstrates the importance of each component in our proposed architecture. We examine the effect of removing excitation signals, harmonic processing, and the impact of using an encoder versus using conditional loudness features. Each configuration highlights different aspects of timbral quality and instrument-specific characteristics.

Ablation Study
Target:
Input
Proposed
No Excitation
No Harmonics
With Encoder