kth.sePublications KTH
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Applying Textual Inversion to Control and Personalize Text-to-Music Models
KTH, School of Electrical Engineering and Computer Science (EECS), Intelligent systems, Speech, Music and Hearing, TMH. KTH Royal Institute of Technology, Stockholm, Sweden.ORCID iD: 0000-0002-8225-5191
KTH, School of Electrical Engineering and Computer Science (EECS), Intelligent systems, Speech, Music and Hearing, TMH.ORCID iD: 0000-0003-2549-6367
KTH, School of Electrical Engineering and Computer Science (EECS), Intelligent systems, Speech, Music and Hearing, TMH.ORCID iD: 0009-0008-7717-2261
KTH, School of Electrical Engineering and Computer Science (EECS), Intelligent systems, Speech, Music and Hearing, TMH.ORCID iD: 0009-0003-8553-3542
2026 (English)In: Machine Learning and Principles and Practice of Knowledge Discovery in Databases - International Workshops of ECML PKDD 2024, Revised Selected Papers, Springer Nature , 2026, Vol. 2559 CCIS, p. 395-401Conference paper, Published paper (Refereed)
Abstract [en]

A text-to-music (TTM) model should synthesize audio that reflects the concepts in a given prompt as long as it has been trained on those concepts. If a prompt references concepts that the TTM model has not been trained on then the audio it synthesizes will likely not match. This paper investigates the application of a simple gradient-based approach called textual inversion (TI) to expand the concept vocabulary of a trained TTM model without compromising the fidelity of concepts on which it has already been trained. We apply this technique to MusicGen and measure its reconstruction and editability quality, as well as its subjective quality. We see TI can expand the concept vocabulary of a pretrained TTM model, thus making it personalized and more controllable without having to finetune the entire model.

Place, publisher, year, edition, pages
Springer Nature , 2026. Vol. 2559 CCIS, p. 395-401
Keywords [en]
Text-to-music, Textual inversion, audio reference
National Category
Other Electrical Engineering, Electronic Engineering, Information Engineering
Identifiers
URN: urn:nbn:se:kth:diva-383413DOI: 10.1007/978-3-032-25305-7_31Scopus ID: 2-s2.0-105040209409OAI: oai:DiVA.org:kth-383413DiVA, id: diva2:2077520
Conference
24th Joint European Conference on Machine Learning and Knowledge Discovery in Databases, ECML PKDD 2024, Vilnius, LTU, Sep 09 2024 - Sep 13 2024
Note

Part of ISBN 9783032253040

QC 20260623

Available from: 2026-06-23 Created: 2026-06-23 Last updated: 2026-06-23Bibliographically approved

Open Access in DiVA

No full text in DiVA

Other links

Publisher's full textScopus

Authority records

Thomé, CarlSturm, BobPertoft, JohnJonason, Nicolas

Search in DiVA

By author/editor
Thomé, CarlSturm, BobPertoft, JohnJonason, Nicolas
By organisation
Speech, Music and Hearing, TMH
Other Electrical Engineering, Electronic Engineering, Information Engineering

Search outside of DiVA

GoogleGoogle Scholar

doi
urn-nbn

Altmetric score

doi
urn-nbn
Total: 8 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf