kth.sePublications KTH
Change search
Link to record
Permanent link

Direct link
Alternative names
Publications (10 of 82) Show all publications
Amerotti, M., Benford, S., Sturm, B. L. .. & Vear, C. (2026). A Live Performance Rule System Informed by Irish Traditional Dance Music. In: Music and Sound Generation in the AI Era - 16th International Symposium, CMMR 2023, Revised Selected Papers: . Paper presented at 16th International Symposium on Computer Music Multidisciplinary Research, CMMR 2023, Tokyo, Japan, November 13-17, 2023 (pp. 127-139). Springer Nature
Open this publication in new window or tab >>A Live Performance Rule System Informed by Irish Traditional Dance Music
2026 (English)In: Music and Sound Generation in the AI Era - 16th International Symposium, CMMR 2023, Revised Selected Papers, Springer Nature , 2026, p. 127-139Conference paper, Published paper (Refereed)
Abstract [en]

This paper describes ongoing work in programming a live performance system for interpreting melodies in ways that mimic Irish traditional dance music practice and that allows plug-and-play human interaction. Existing performance systems are almost exclusively aimed at piano performance and classical music and few are aimed specifically at traditional music. We develop a rule-based approach using expert knowledge that converts a melody into control parameters to synthesize an expressive MIDI performance, focusing on ornamentation, dynamics and subtle time deviation. Furthermore, we make the system controllable (e.g., via knobs or expression pedals) such that it can be controlled in real time by a musician. Our preliminary evaluations show the system can render expressive performances mimicking traditional practice and allows for engaging with Irish traditional dance music in new ways. We provide several examples online (See this website: https://www.kth.se/profile/bobs/page/research-data).

Place, publisher, year, edition, pages
Springer Nature, 2026
Keywords
Irish, music performance modeling, traditional music
National Category
Other Electrical Engineering, Electronic Engineering, Information Engineering Computer Sciences Music Musicology
Identifiers
urn:nbn:se:kth:diva-372795 (URN)10.1007/978-3-032-02042-0_9 (DOI)001660305400009 ()2-s2.0-105020024736 (Scopus ID)
Conference
16th International Symposium on Computer Music Multidisciplinary Research, CMMR 2023, Tokyo, Japan, November 13-17, 2023
Note

Not duplicate with diva 1795616 

Part of ISBN 9783032020413

QC 20251118

Available from: 2025-11-18 Created: 2025-11-18 Last updated: 2026-05-29Bibliographically approved
Thomé, C., Sturm, B., Pertoft, J. & Jonason, N. (2026). Applying Textual Inversion to Control and Personalize Text-to-Music Models. In: Machine Learning and Principles and Practice of Knowledge Discovery in Databases - International Workshops of ECML PKDD 2024, Revised Selected Papers: . Paper presented at 24th Joint European Conference on Machine Learning and Knowledge Discovery in Databases, ECML PKDD 2024, Vilnius, LTU, Sep 09 2024 - Sep 13 2024 (pp. 395-401). Springer Nature, 2559 CCIS
Open this publication in new window or tab >>Applying Textual Inversion to Control and Personalize Text-to-Music Models
2026 (English)In: Machine Learning and Principles and Practice of Knowledge Discovery in Databases - International Workshops of ECML PKDD 2024, Revised Selected Papers, Springer Nature , 2026, Vol. 2559 CCIS, p. 395-401Conference paper, Published paper (Refereed)
Abstract [en]

A text-to-music (TTM) model should synthesize audio that reflects the concepts in a given prompt as long as it has been trained on those concepts. If a prompt references concepts that the TTM model has not been trained on then the audio it synthesizes will likely not match. This paper investigates the application of a simple gradient-based approach called textual inversion (TI) to expand the concept vocabulary of a trained TTM model without compromising the fidelity of concepts on which it has already been trained. We apply this technique to MusicGen and measure its reconstruction and editability quality, as well as its subjective quality. We see TI can expand the concept vocabulary of a pretrained TTM model, thus making it personalized and more controllable without having to finetune the entire model.

Place, publisher, year, edition, pages
Springer Nature, 2026
Keywords
Text-to-music, Textual inversion, audio reference
National Category
Other Electrical Engineering, Electronic Engineering, Information Engineering
Identifiers
urn:nbn:se:kth:diva-383413 (URN)10.1007/978-3-032-25305-7_31 (DOI)2-s2.0-105040209409 (Scopus ID)
Conference
24th Joint European Conference on Machine Learning and Knowledge Discovery in Databases, ECML PKDD 2024, Vilnius, LTU, Sep 09 2024 - Sep 13 2024
Note

Part of ISBN 9783032253040

QC 20260623

Available from: 2026-06-23 Created: 2026-06-23 Last updated: 2026-06-23Bibliographically approved
Amerotti, M., Benford, S., Hazzard, A., Martinez Avila, J., Tennent, P. & Sturm, B. (2026). Charting Creative Journeys through Musical Performance with AI. In: C and C 2026 - Proceedings of the 2026 Conference on Creativity and Cognition: . Paper presented at 18th ACM Creativity and Cognition Conference, C and C 2026, London, United Kingdom, Jul 13-16 2026 (pp. 911-920). Association for Computing Machinery (ACM)
Open this publication in new window or tab >>Charting Creative Journeys through Musical Performance with AI
Show others...
2026 (English)In: C and C 2026 - Proceedings of the 2026 Conference on Creativity and Cognition, Association for Computing Machinery (ACM) , 2026, p. 911-920Conference paper, Published paper (Refereed)
Abstract [en]

We chart the creative journeys of two groups of musicians as they explored how to engage with an AI system developed for the live performance of Irish Traditional Dance Music. A group of expert Irish musicians participated in an extended workshop to refine the system while a professional contemporary folk duo composed and rehearsed material for a public show. Drawing on the standard definition of creativity, we chart their creative journeys across a landscape we define by the two dimensions of originality and effectiveness. We reveal the importance of authenticity to a traditional practice as an opposing force to creativity and consider how the tension between these shaped their musical journeys. We generalise five strategies that move creative practitioners across this landscape in different directions: proving your chops, finding your place, perfecting the nuances, innovating, and glitching.

Place, publisher, year, edition, pages
Association for Computing Machinery (ACM), 2026
Keywords
AI, Irish traditional music, authenticity, creativity, journeys, trajectories
National Category
Music Musicology
Identifiers
urn:nbn:se:kth:diva-387005 (URN)10.1145/3803784.3807527 (DOI)2-s2.0-105045820369 (Scopus ID)
Conference
18th ACM Creativity and Cognition Conference, C and C 2026, London, United Kingdom, Jul 13-16 2026
Note

Part of ISBN 9798400725838

QC 20260812

Available from: 2026-08-12 Created: 2026-08-12 Last updated: 2026-08-12Bibliographically approved
Dalmazzo, D., Déguernel, K. & Sturm, B. (2026). ChromaFlow: Modeling And Generating Harmonic Progressions With a Transformer And Voicing Encoding. In: Machine Learning and Principles and Practice of Knowledge Discovery in Databases - International Workshops of ECML PKDD 2024, Revised Selected Papers: . Paper presented at 24th Joint European Conference on Machine Learning and Knowledge Discovery in Databases, ECML PKDD 2024, Vilnius, LTU, Sep 09 2024 - Sep 13 2024 (pp. 354-363). Springer Nature, 2559 CCIS
Open this publication in new window or tab >>ChromaFlow: Modeling And Generating Harmonic Progressions With a Transformer And Voicing Encoding
2026 (English)In: Machine Learning and Principles and Practice of Knowledge Discovery in Databases - International Workshops of ECML PKDD 2024, Revised Selected Papers, Springer Nature , 2026, Vol. 2559 CCIS, p. 354-363Conference paper, Published paper (Refereed)
Abstract [en]

Modeling harmonic progressions in symbolic music is a complex task that requires generating musically coherent and varied chord sequences. In this study, we employ a transformer-based architecture trained on a comprehensive dataset of 48,072 songs, which includes an augmented set of 4,300 original pieces from the iReal Pro application transposed across all chromatic keys. We introduce a novel tokenization and voicing encoding strategy designed to enhance the musicality of the generated chord progressions. Our approach not only generates chord progression suggestions but also provides corresponding voicings tailored for instruments such as piano and guitar. To evaluate the effectiveness of our model, we conducted a listening test comparing the harmonic progressions produced by our approach against those from a baseline model. The results indicate that our model generates progressions with more fluid voicings, coherent harmonic motion, and plausible chord suggestions, effectively utilizing repetition and variation to enhance musicality.

Place, publisher, year, edition, pages
Springer Nature, 2026
Keywords
Chord Progressions, Music Generation, The Transformer Network
National Category
Natural Language Processing
Identifiers
urn:nbn:se:kth:diva-383417 (URN)10.1007/978-3-032-25305-7_26 (DOI)2-s2.0-105040262200 (Scopus ID)
Conference
24th Joint European Conference on Machine Learning and Knowledge Discovery in Databases, ECML PKDD 2024, Vilnius, LTU, Sep 09 2024 - Sep 13 2024
Note

Part of ISBN 9783032253040

QC 20260612

Available from: 2026-06-12 Created: 2026-06-12 Last updated: 2026-06-12Bibliographically approved
Casini, L., Cros Vila, L., Dalmazzo, D., Kaila, A.-K. & Sturm, B. L. .. (2026). Data‑Driven Analysis of Text‑Conditioning in AI‑Generated Music: A Case Study with Suno and Udio. Transactions of the International Society for Music Information Retrieval, 9(1), 194-209
Open this publication in new window or tab >>Data‑Driven Analysis of Text‑Conditioning in AI‑Generated Music: A Case Study with Suno and Udio
Show others...
2026 (English)In: Transactions of the International Society for Music Information Retrieval, ISSN 2514-3298, Vol. 9, no 1, p. 194-209Article in journal (Refereed) Published
Abstract [en]

Online commercial artificial intelligence (AI) platforms for generating music from text prompts (AI music) are now being used by many users to create millions of music audio recordings daily. Some AI music is appearing in advertising, music playlists of restaurants and gyms, and even hit music charts, in many countries. How are users engaging with these text‑to‑music AI platforms, where text is a principal mode of interaction to specify prompts (e.g., free terms), lyrics (e.g., sung terms), and tags (e.g., high‑level stylistic terms)? What languages appear? What characterizes prompts, lyrics, and tags? How are mentions of real artists used? What kind of additional instructions (metatags) are used? To address these questions, we assemble and analyze a collection of 101, 953 songs generated from May to October 2024 by 60, 342 users of Suno and Udio. Using a combination of state‑of‑the‑art text‑embedding models, dimensionality reduction, and clustering methods, we analyze the prompts, tags, and lyrics and automatically annotate and display the processed data in interactive plots. Our results reveal prominent themes in lyrics, language preferences, and prompting strategies, as well as peculiar attempts at steering models through the use of metatags. We share our code and data resources to promote further musicological study of AI music.

Place, publisher, year, edition, pages
Ubiquity Press, Ltd., 2026
Keywords
AI music, generative AI, Suno, Udio, exploratory data analysis, natural language processing
National Category
Musicology Artificial Intelligence
Research subject
Media Technology; Computer Science
Identifiers
urn:nbn:se:kth:diva-381510 (URN)10.5334/tismir.273 (DOI)
Note

QC 20260522

Available from: 2026-05-18 Created: 2026-05-18 Last updated: 2026-05-22Bibliographically approved
Sturm, B. & Kanhov, E. (2026). oljud—ʇᴉnɹq(n): A Sonic Manifesto of Resistance to Generative AI in Music. In: International Computer Music Conference 2026 May 10-16, 2026  Hamburg, Germany: . Paper presented at International Computer Music Conference 2026 May 10-16, 2026, Hamburg, Germany (pp. 352-359).
Open this publication in new window or tab >>oljud—ʇᴉnɹq(n): A Sonic Manifesto of Resistance to Generative AI in Music
2026 (English)In: International Computer Music Conference 2026 May 10-16, 2026  Hamburg, Germany, 2026, p. 352-359Conference paper, Published paper (Refereed)
Abstract [en]

Some discourses about music and AI are frustratingly shallow and insular, raising outdated musical tropes and ignoring modern developments, flattening the rich and varied functions of music in life, and overlooking serious ethical issues with the technology (creating it, maintaining it, using it, and imposing it). We respond to this shallowness via artistic activism and agonistic artistic research, resulting in the site-specific work “oljud—bruit(n)”. Our action was directed at an “AI music composition” competition in2025 organised as part of a very expensive engineering workshop focused on music generation research, but our motivations are more broad. This paper records the context, composition and realization of our “sonic manifesto of resistance”, which was ultimately disqualified from the competition.

National Category
Music
Identifiers
urn:nbn:se:kth:diva-383075 (URN)
Conference
International Computer Music Conference 2026 May 10-16, 2026, Hamburg, Germany
Funder
Swedish Research Council, 2024-01012
Note

This is an open-access article distributed under the terms of the Creative Commons Attribution License 3.0 Unported, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.

QC 20260605

Available from: 2026-06-05 Created: 2026-06-05 Last updated: 2026-08-11Bibliographically approved
Dalmazzo, D., Déguernel, K. & Sturm, B. (2025). A Computer Application to Explore 53-Tone Equal Temperament Harmonies Through Modal Interchange. In: Proceedings of the International Conference on New Interfaces for Musical Expression: . Paper presented at 25th International Conference on New Interfaces for Musical Expression, NIME 2025, Canberra, Australia, Jun 24 2025 - Jun 27 2025. International Conference on New Interfaces for Musical Expression
Open this publication in new window or tab >>A Computer Application to Explore 53-Tone Equal Temperament Harmonies Through Modal Interchange
2025 (English)In: Proceedings of the International Conference on New Interfaces for Musical Expression, International Conference on New Interfaces for Musical Expression , 2025Conference paper, Published paper (Refereed)
Abstract [en]

We present a novel computer application for real-time exploration of microtonal harmonies through intuitive visualization and integrated MIDI controllers, bridging theoretical concepts and practical musical applications. We extend modern harmonic principles from 12-tone equal temperament (12-TET), incorporating interval distinctions from 31-TET (subminor, neutral, supermajor), and further expanding into detailed harmonic possibilities of 53-tone equal temperament (53-TET). Our application leverages modal interchange and parallel chord substitutions, offering intuitive navigation through microtonal harmonic trajectories. The implementation utilizes MIDI Polyphonic Expression (MPE) via our custom MaxForLive application, The Bridge, ensuring precise microtonal control and compatibility with digital audio workstations (Ableton Live 11/12). The system includes real-time visualization, interactive chord manipulation, and a comprehensive Scale Editor for harmonic experimentation. Through practical examples and theoretical analysis, we demonstrate how this approach reveals new harmonic possibilities while maintaining connections to established modal frameworks. This research contributes to the growing microtonal music field by providing theoretical foundations and practical tools that incorporate extended tuning systems into contemporary musical practice.

Place, publisher, year, edition, pages
International Conference on New Interfaces for Musical Expression, 2025
Keywords
53 Tone Equal Temperament, Data Visualization, Harmonic Exploration, Microtonality, Modal Interchange, Music Theory
National Category
Musicology
Identifiers
urn:nbn:se:kth:diva-385684 (URN)10.5281/zenodo.15698843 (DOI)2-s2.0-105010896111 (Scopus ID)
Conference
25th International Conference on New Interfaces for Musical Expression, NIME 2025, Canberra, Australia, Jun 24 2025 - Jun 27 2025
Note

QC 20260721

Available from: 2026-07-21 Created: 2026-07-21 Last updated: 2026-07-21Bibliographically approved
Sturm, B. & Kanhov, E. (2025). En Svår Jul: An Album of Unpracticed, Unpolished and Unproduced Christmas Music.
Open this publication in new window or tab >>En Svår Jul: An Album of Unpracticed, Unpolished and Unproduced Christmas Music
2025 (English)Artistic output (Unrefereed)
National Category
Music
Identifiers
urn:nbn:se:kth:diva-373317 (URN)
Note

QC 20260529

Available from: 2025-11-28 Created: 2025-11-28 Last updated: 2026-05-29Bibliographically approved
Grouwels, J., Jonason, N. & Sturm, B. (2025). Exploring the Expressive Space of an Articulatory Vocal Modal using Quality-Diversity Optimization with Multimodal Embeddings. In: GECCO 2025 - Proceedings of the 2025 Genetic and Evolutionary Computation Conference: . Paper presented at 2025 Genetic and Evolutionary Computation Conference, GECCO 2025, Malaga, Spain, Jul 14 2025 - Jul 18 2025 (pp. 1362-1370). Association for Computing Machinery (ACM)
Open this publication in new window or tab >>Exploring the Expressive Space of an Articulatory Vocal Modal using Quality-Diversity Optimization with Multimodal Embeddings
2025 (English)In: GECCO 2025 - Proceedings of the 2025 Genetic and Evolutionary Computation Conference, Association for Computing Machinery (ACM) , 2025, p. 1362-1370Conference paper, Published paper (Refereed)
Abstract [en]

Knowing which sounds can be produced by a simulated vocal model and how they are connected to its articulatory behavior is not trivial. Being able to map this out can be interesting for applications that make use of the extended capabilities of a voice, e.g., singing or vocal imitations. We present a method that achieves this for a state-of-the-art articulatory vocal model (VocalTractLab) by combining it with a recent Quality-Diversity algorithm (CMA-MAE) and audio embeddings obtained through a multi-modal pretrained model (CLAP). The text-capabilities of CLAP make it possible to steer the exploration through a text prompt. We show that the method explores more efficiently than a random sampling baseline, covering more of the measure space and achieving higher objective scores. We provide several listening examples and the source code for a scalable implementation.

Place, publisher, year, edition, pages
Association for Computing Machinery (ACM), 2025
Keywords
articulatory vocal model, CLAP, CMA-MAE, diversity optimization, multimodal, quality-diversity, text prompt, VocalTractLab
National Category
Natural Language Processing Computer Sciences Comparative Language Studies and Linguistics Signal Processing
Identifiers
urn:nbn:se:kth:diva-369365 (URN)10.1145/3712256.3726313 (DOI)001556459900153 ()2-s2.0-105013082602 (Scopus ID)
Conference
2025 Genetic and Evolutionary Computation Conference, GECCO 2025, Malaga, Spain, Jul 14 2025 - Jul 18 2025
Note

Part of ISBN 9798400714658

QC 20250903

Available from: 2025-09-03 Created: 2025-09-03 Last updated: 2025-12-05Bibliographically approved
Huang, R. S., Holzapfel, A. & Sturm, B. L. T. (2025). FROM PHILOSOPHY TO PRACTICE: A Culturally Informed Ethics of Music AI in Asia (2ed.). In: Artificial Intelligence and Music Ecosystem: Second Edition (pp. 171-187). Taylor and Francis
Open this publication in new window or tab >>FROM PHILOSOPHY TO PRACTICE: A Culturally Informed Ethics of Music AI in Asia
2025 (English)In: Artificial Intelligence and Music Ecosystem: Second Edition, Taylor and Francis , 2025, 2, p. 171-187Chapter in book (Other academic)
Place, publisher, year, edition, pages
Taylor and Francis, 2025 Edition: 2
National Category
Musicology Pedagogy
Identifiers
urn:nbn:se:kth:diva-383082 (URN)2-s2.0-105039971777 (Scopus ID)
Note

Part of ISBN 9781040436844, 9781032846262

QC 20260608

Available from: 2026-06-08 Created: 2026-06-08 Last updated: 2026-06-08Bibliographically approved
Organisations
Identifiers
ORCID iD: ORCID iD iconorcid.org/0000-0003-2549-6367

Search in DiVA

Show all publications