kth.sePublications KTH
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Neural music instrument cloning from few samples
KTH, School of Electrical Engineering and Computer Science (EECS), Intelligent systems, Speech, Music and Hearing, TMH.
KTH, School of Electrical Engineering and Computer Science (EECS), Intelligent systems, Speech, Music and Hearing, TMH.ORCID iD: 0000-0003-2549-6367
2022 (English)In: Proceedings of the 25th International Conference on Digital Audio Effects (DAFx20in22), 2022, Vol. 3, p. 296-303Conference paper, Published paper (Refereed)
Abstract [en]

Neural music instrument cloning is an application of deep neural networks for imitating the timbre of a particular music instrument recording with a trained neural network. One can create suchclones using an approach such as DDSP, which has been shownto achieve good synthesis quality for several instrument types.However, this approach needs about ten minutes of audio datafrom the instrument of interest (target recording audio). In thiswork, we modify the DDSP architecture and apply transfer learning techniques used in speech voice cloning to significantlyreduce the amount of target recording audio required. We compare various cloning approaches and architectures across durationsof target recording audio, ranging from four to 256 seconds. Wedemonstrate editing of loudness and pitch as well as timbre transfer from only 16 seconds of target recording audio. Our code isavailable online1as well as many audio examples.

Place, publisher, year, edition, pages
2022. Vol. 3, p. 296-303
National Category
Signal Processing
Identifiers
URN: urn:nbn:se:kth:diva-326017Scopus ID: 2-s2.0-85138790240OAI: oai:DiVA.org:kth-326017DiVA, id: diva2:1752391
Conference
25th International Conference on Digital Audio Effects (DAFx20in22), Vienna, Austria, September 2022
Funder
EU, Horizon 2020, 864189
Note

QC 20230424

Available from: 2023-04-21 Created: 2023-04-21 Last updated: 2026-03-05Bibliographically approved
In thesis
1. Bittersweet Lessons in Music AI Research: Neural Instrument Synthesis, Multi-modal Representations, Symbolic Music Generation
Open this publication in new window or tab >>Bittersweet Lessons in Music AI Research: Neural Instrument Synthesis, Multi-modal Representations, Symbolic Music Generation
2026 (English)Doctoral thesis, comprehensive summary (Other academic)
Abstract [en]

This compilation thesis explores AI techniques in three areas related to music making: neural instrument synthesis, multi-modal representations, and symbolic music generation. In neural instrument synthesis, we explore architectural changes and transfer learning techniques to apply neural synthesis methods to instruments where little data is available. We then move to zero-shot audio applications of multi-modal representations, including text-guided audio equalization, visualization of instrument sounds, and text-driven synthesizer programming. In the domain of symbolic music, we propose superposed language modelling, a generalisation of masked language modelling that enables controllable generation and editing of music using event-attribute domain constraints. We then experiment with text-driven music generation and editing with LLMs augmented with a retrieval system to fetch relevant few-shot examples, showing early signs that LLMs could challenge domain specific approaches to symbolic music generation. We then bridge the symbolic and audio domains by using an audio-domain model of human preferences as a reward to tune a symbolic music generation model, producing music which according to the preference model is better than Mozart. Reflecting on our work, we focus on data availability as the key factor in determining Music AI capabilities and that much of our work in this thesis can be seen as capability arbitrage: redirecting capabilities from data-rich domains towards data-poor domains. We conclude by speculating on music making capabilities of future AI considering the massive iceberg of data that remains unused.

Abstract [sv]

Denna avhandling utforskar AI-tekniker inom tre områden relaterade till musikskapande: neural instrumentsyntes, multimodala representationer och symbolisk musikgenerering. Inom neural instrumentsyntes utforskar vi arkitekturförändringar och överföringsinlärning för att tillämpa neurala syntesmetoder på instrument där lite data finns tillgänglig. Vi övergår sedan till zero-shot-ljudtillämpningar av multimodala representationer, inklusive textguidad ljudekvalisering, visualisering av instrumentljud och textdriven synthesizerprogrammering. Inom symbolisk musik föreslår vi superponerade språkmodeller, en generalisering av maskerade språkmodeller för kontrollerbar generering och redigering av musik med event-attribut-domänbegränsningar. Vi experimenterar sedan med textdriven musikgenerering och redigering med LLM:er förstärkta med ett retrieval-system för att hämta relevanta few-shot-exempel, ett tidigt tecken på att LLM:er kan utmana domänspecifika metoder för symbolisk musikgenerering. Vi överbryggar sedan de symboliska och ljuddomänerna genom att använda en ljuddomänmodell av mänskliga preferenser som belöningssignal för att finjustera en symbolisk musikgenereringsmodell, och producerar musik som enligt preferensmodellen är bättre än Mozart. I en reflektion kring vårt arbete lyfter vi datatillgänglighet som den avgörande faktorn för musik-AI:s förmågor, och att mycket av vårt arbete kan ses som capability arbitrage: en omdirigering av förmågor från datarika domäner mot datafattiga domäner. Vi avslutar med att spekulera kring framtida AI-förmågor för musikskapande med hänsyn till det massiva isberg av data som fortfarande inte nyttjas.

Place, publisher, year, edition, pages
Stockholm: KTH Royal Institute of Technology, 2026. p. xvi, 67
Series
TRITA-EECS-AVL ; 2026:19
Keywords
Artificial Intelligence, Machine Learning, Music
National Category
Artificial Intelligence
Research subject
Speech and Music Communication
Identifiers
urn:nbn:se:kth:diva-377801 (URN)978-91-8106-542-8 (ISBN)
Public defence
2026-03-27, https://kth-se.zoom.us/j/64932870406, F3 (Flodis), Lindstedtsvägen 26 & 28, Stockholm, Sweden, 15:00 (English)
Opponent
Supervisors
Funder
EU, Horizon 2020, 864189
Note

QC 20260306

Available from: 2026-03-06 Created: 2026-03-05 Last updated: 2026-03-30Bibliographically approved

Open Access in DiVA

fulltext(785 kB)435 downloads
File information
File name FULLTEXT01.pdfFile size 785 kBChecksum SHA-512
68e243a7cfd956dd64d310515567b01fa8ef01d44b31ad3f7a8c30430e99c836860a496ee060bfb50f414aece70d0f6212ce48396dbc2d43604b9ea568c383c9
Type fulltextMimetype application/pdf

Other links

ScopusConference websiteConference proceedings

Authority records

Jonason, NicolasSturm, Bob

Search in DiVA

By author/editor
Jonason, NicolasSturm, Bob
By organisation
Speech, Music and Hearing, TMH
Signal Processing

Search outside of DiVA

GoogleGoogle Scholar
Total: 438 downloads
The number of downloads is the sum of all downloads of full texts. It may include eg previous versions that are now no longer available

urn-nbn

Altmetric score

urn-nbn
Total: 1127 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf