kth.sePublications KTH
Change search
Link to record
Permanent link

Direct link
Gustafsson, Joakim, ProfessorORCID iD iconorcid.org/0000-0002-0397-6442
Alternative names
Publications (10 of 178) Show all publications
Kontogiorgos, D., Gustafsson, J. & Shah, J. (2026). Automatic Early Detection of Explanation Needs in Human–Robot Interaction. In: : . Paper presented at The 21st ACM/IEEE International Conference on Human-Robot Interaction, Edinburgh, Scotland, UK, March 16–19, 2026.
Open this publication in new window or tab >>Automatic Early Detection of Explanation Needs in Human–Robot Interaction
2026 (English)Conference paper, Published paper (Refereed)
Abstract [en]

Enabling robots to display the reasoning behind decisions requires them to detect when explanations are needed by users. A crucial driver of explanation need is that it often manifests implicitly: users exhibit behavioural signals indicating misalignment well before they explicitly request an explanation. Psychological studies show that in human interactions, such needs are sensed through multimodal cues and addressed through the co-construction of explanations in real time. Building on this, we introduce an approach for the early detection of explanation needs in HRI. Our method recognises when an explanation is likely to become necessary, enabling robots to act proactively. We evaluate the approach on an existing HRI dataset using features describing facial expressions,body movement, and vocal behaviour, combined with time-series classification techniques. Our results show that different classes of learning algorithms (unsupervised anomaly-based methods and supervised classification models) offer complementary strengths for detecting explanation needs. In particular, unsupervised methods enable early warning signals when labels are unavailable, while supervised models provide stronger discrimination (AUROC 0.7)when annotated data is available. We discuss the implications of these findings for the development of explanation-capable robots and outline future directions for proactive explanations in HRI.

Keywords
explainability, social signal processing, multimodality
National Category
Human Computer Interaction
Identifiers
urn:nbn:se:kth:diva-378471 (URN)
Conference
The 21st ACM/IEEE International Conference on Human-Robot Interaction, Edinburgh, Scotland, UK, March 16–19, 2026
Available from: 2026-03-21 Created: 2026-03-21 Last updated: 2026-03-23
Esteve, K., Fredriksson, M., Gustafsson, J., Kontogiorgos, D. & Mashiyi-Veikkola, T. (2026). Towards a Proactive Cooking Companion for the Elderly. In: IWSDS 2026 - 16th International Workshop on Spoken Dialogue Systems Technology, Proceedings of the Conference: . Paper presented at 16th International Workshop on Spoken Dialogue Systems Technology, IWSDS 2026, Trento, Italy, Feb 26 2026 - Mar 01 2026 (pp. 134-141). Association for Computational Linguistics (ACL)
Open this publication in new window or tab >>Towards a Proactive Cooking Companion for the Elderly
Show others...
2026 (English)In: IWSDS 2026 - 16th International Workshop on Spoken Dialogue Systems Technology, Proceedings of the Conference, Association for Computational Linguistics (ACL) , 2026, p. 134-141Conference paper, Published paper (Refereed)
Abstract [en]

We present a voice assistant designed as a cooking companion, addressing both nutritional and social needs through intelligent interaction. Through WoZ experiments, we validated: social dialogue serves functional purposes, where “chatty” assistants transform cooking pauses into engaging interactions while instructional-only versions create frustrating dead air, despite identical timing.

Place, publisher, year, edition, pages
Association for Computational Linguistics (ACL), 2026
National Category
Human Computer Interaction Computer Sciences
Identifiers
urn:nbn:se:kth:diva-383415 (URN)2-s2.0-105040220861 (Scopus ID)
Conference
16th International Workshop on Spoken Dialogue Systems Technology, IWSDS 2026, Trento, Italy, Feb 26 2026 - Mar 01 2026
Note

Part of ISBN 9781952148255

QC 20260617

Available from: 2026-06-17 Created: 2026-06-17 Last updated: 2026-06-18Bibliographically approved
Marcinek, L., Beskow, J. & Gustafsson, J. (2025). A Dual-Control Dialogue Framework for Human-Robot Interaction Data Collection: Integrating Human Emotional and Contextual Awareness with Conversational AI. In: Social Robotics - 16th International Conference, ICSR + AI 2024, Proceedings: . Paper presented at 16th International Conference on Social Robotics, ICSR + AI 2024, Odense, Denmark, Oct 23 2024 - Oct 26 2024 (pp. 290-297). Springer Nature
Open this publication in new window or tab >>A Dual-Control Dialogue Framework for Human-Robot Interaction Data Collection: Integrating Human Emotional and Contextual Awareness with Conversational AI
2025 (English)In: Social Robotics - 16th International Conference, ICSR + AI 2024, Proceedings, Springer Nature , 2025, p. 290-297Conference paper, Published paper (Refereed)
Abstract [en]

This paper presents a dialogue framework designed to capture human-robot interactions enriched with human-level situational awareness. The system integrates advanced large language models with real-time human-in-the-loop control. Central to this framework is an interaction manager that oversees information flow, turn-taking, and prosody control of a social robot’s responses. A key innovation is the control interface, enabling a human operator to perform tasks such as emotion recognition and action detection through a live video feed. The operator also manages high-level tasks, like topic shifts or behaviour instructions. Input from the operator is incorporated into the dialogue context managed by GPT-4o, thereby influencing the ongoing interaction. This allows for the collection of interactional data from an automated system that leverages human-level emotional and situational awareness. The audio-visual data will be used to explore the impact of situational awareness on user behaviors in task-oriented human-robot interaction.

Place, publisher, year, edition, pages
Springer Nature, 2025
Keywords
Dialogue system, Emotions, Situational Context
National Category
Natural Language Processing Human Computer Interaction Computer Sciences
Identifiers
urn:nbn:se:kth:diva-362497 (URN)10.1007/978-981-96-3519-1_27 (DOI)001531735400027 ()2-s2.0-105002141806 (Scopus ID)
Conference
16th International Conference on Social Robotics, ICSR + AI 2024, Odense, Denmark, Oct 23 2024 - Oct 26 2024
Note

Part of ISBN 9789819635184

QC 20250424

Available from: 2025-04-16 Created: 2025-04-16 Last updated: 2025-12-08Bibliographically approved
Francis, J., Gustafsson, J. & Székely, É. (2025). From Static to Dynamic: Enhancing AAC with Generative Imagery and Zero-Shot TTS. In: Interspeech 2025: . Paper presented at 26th Interspeech Conference 2025, Rotterdam, Netherlands, Kingdom of the, August 17-21, 2025 (pp. 4960-4962). International Speech Communication Association
Open this publication in new window or tab >>From Static to Dynamic: Enhancing AAC with Generative Imagery and Zero-Shot TTS
2025 (English)In: Interspeech 2025, International Speech Communication Association , 2025, p. 4960-4962Conference paper, Published paper (Refereed)
Abstract [en]

This paper presents an Augmentative and Alternative Communication (AAC) approach for minimally verbal children with Autism Spectrum Disorder. Traditional AAC systems use fixed symbol sets and pre-defined Text-to-Speech (TTS) voices, this proposed method leverages text-to-image generation and zero-shot TTS to expand expressive capabilities. Users can create visual symbols for concepts and interests, enabling richer communication. Further, zero-shot TTS allows users to upload or record personalized voices, enabling users to have individualized output. By minimizing reliance on static symbols and voices, this approach aims to increase communicative agency, personal relevance, and social validity, areas often neglected in traditional interventions. Future research will explore long-term effects on communicative skills, user satisfaction, social engagement, and adaptability across various cultural and linguistic settings, aiming to develop more dynamic and personalized AAC solutions.

Place, publisher, year, edition, pages
International Speech Communication Association, 2025
Keywords
AAC, Human-Computer Interaction, Speech Synthesis, TTS
National Category
Natural Language Processing Human Computer Interaction Other Engineering and Technologies
Identifiers
urn:nbn:se:kth:diva-372783 (URN)10.21437/Interspeech.2025-2815 (DOI)001613931400417 ()2-s2.0-105020070493 (Scopus ID)
Conference
26th Interspeech Conference 2025, Rotterdam, Netherlands, Kingdom of the, August 17-21, 2025
Note

QC 20251124

Available from: 2025-11-24 Created: 2025-11-24 Last updated: 2026-05-29Bibliographically approved
Marcinek, L., Irfan, B., Skantze, G., Abelho Pereira, A. T. & Gustafsson, J. (2025). Role of Reasoning in LLM Enjoyment Detection: Evaluation Across Conversational Levels for Human-Robot Interaction. In: Frédéric Béchet, Fabrice Lefèvre, Nicholas Asher, Seokhwan Kim, Teva Merlin (Ed.), Proceedings of the 26th Annual Meeting of the Special Interest Group on Discourse and Dialogue: . Paper presented at The 26th Annual Meeting of the Special Interest Group on Discourse and Dialogue, Avignon, France, Aug 25-27, 2025 (pp. 573-590). ASSOC COMPUTATIONAL LINGUISTICS
Open this publication in new window or tab >>Role of Reasoning in LLM Enjoyment Detection: Evaluation Across Conversational Levels for Human-Robot Interaction
Show others...
2025 (English)In: Proceedings of the 26th Annual Meeting of the Special Interest Group on Discourse and Dialogue / [ed] Frédéric Béchet, Fabrice Lefèvre, Nicholas Asher, Seokhwan Kim, Teva Merlin, ASSOC COMPUTATIONAL LINGUISTICS , 2025, p. 573-590Conference paper, Published paper (Refereed)
Abstract [en]

User enjoyment is central to developing conversational AI systems that can recover from failures and maintain interest over time. However, existing approaches often struggle to detect subtle cues that reflect user experience. Large Language Models (LLMs) with reasoning capabilities have outperformed standard models on various other tasks, suggesting potential benefits for enjoyment detection. This study investigates whether models with reasoning capabilities outperform standard models when assessing enjoyment in a human-robot dialogue corpus at both turn and interaction levels. Results indicate that reasoning capabilities have complex, model-dependent effects rather than universal benefits. While performance was nearly identical at the interaction level (0.44 vs 0.43), reasoning models substantially outperformed at the turn level (0.42 vs 0.36). Notably, LLMs correlated better with users’ self-reported enjoyment metrics than human annotators, despite achieving lower accuracy against human consensus ratings. Analysis revealed distinctive error patterns: non-reasoning models showed bias toward positive ratings at the turn level, while both model types exhibited central tendency bias at the interaction level. These findings suggest that reasoning should be applied selectively based on model architecture and assessment context, with assessment granularity significantly influencing relative effectiveness.

Place, publisher, year, edition, pages
ASSOC COMPUTATIONAL LINGUISTICS, 2025
National Category
Computer and Information Sciences Natural Language Processing Computer graphics and computer vision
Identifiers
urn:nbn:se:kth:diva-374881 (URN)001734552600046 ()
Conference
The 26th Annual Meeting of the Special Interest Group on Discourse and Dialogue, Avignon, France, Aug 25-27, 2025
Note

Part of ISBN 979-8-89176-329-6

QC 20260107

Available from: 2026-01-06 Created: 2026-01-06 Last updated: 2026-06-09Bibliographically approved
Marcinek, L., Beskow, J. & Gustafsson, J. (2025). Towards Adaptable and Intelligible Speech Synthesis in Noisy Environments. In: Interspeech 2025: . Paper presented at 26th Interspeech Conference 2025, Rotterdam, Netherlands, Kingdom of the, August 17-21, 2025 (pp. 2165-2169). International Speech Communication Association
Open this publication in new window or tab >>Towards Adaptable and Intelligible Speech Synthesis in Noisy Environments
2025 (English)In: Interspeech 2025, International Speech Communication Association , 2025, p. 2165-2169Conference paper, Published paper (Refereed)
Abstract [en]

We present an investigation into adaptable speech synthesis for noisy environments. Leveraging a zero-shot TTS we synthesized a corpus of 1,200 speech samples from 100 sentences of varying complexity, each generated at six distinct levels of vocal effort. To simulate realistic listening conditions, the synthesized speech is merged with environmental noise recordings from a diverse range of indoor and transportation settings at nine different signal-to-noise ratios. We assess the intelligibility of the resulting noisy speech using the ASR word error rates across conditions. Additionally, the input text was evaluated using four metrics on sentence complexity and word predictability. A number of regression models that used noise type, SNR, vocal effort and text as input were trained to predict ASR WER. Results show that increased vocal effort improves intelligibility, with benefits up to 30% in adverse conditions, most most pronounced in environments with competing speech at low SNRs.

Place, publisher, year, edition, pages
International Speech Communication Association, 2025
Keywords
noisy environments, speech adaptation, speech intelligibility, speech synthesis
National Category
Natural Language Processing Signal Processing Computer Sciences
Identifiers
urn:nbn:se:kth:diva-372805 (URN)10.21437/Interspeech.2025-2787 (DOI)001585350500441 ()2-s2.0-105020064005 (Scopus ID)
Conference
26th Interspeech Conference 2025, Rotterdam, Netherlands, Kingdom of the, August 17-21, 2025
Note

QC 20251113

Available from: 2025-11-13 Created: 2025-11-13 Last updated: 2026-05-29Bibliographically approved
Lameris, H., Gustafsson, J. & Székely, É. (2025). VoiceQualityVC: A Voice Conversion System for Studying the Perceptual Effects of Voice Quality in Speech. In: Interspeech 2025: . Paper presented at 26th Interspeech Conference 2025, Rotterdam, Netherlands, Kingdom of the, August 17-21, 2025 (pp. 2295-2299). International Speech Communication Association
Open this publication in new window or tab >>VoiceQualityVC: A Voice Conversion System for Studying the Perceptual Effects of Voice Quality in Speech
2025 (English)In: Interspeech 2025, International Speech Communication Association , 2025, p. 2295-2299Conference paper, Published paper (Refereed)
Abstract [en]

Voice quality is an often overlooked aspect of speech with many communicative functions. Voice quality conveys both paralinguistic and pragmatic information, such as signalling speaker stance and aids in grounding. In this paper, we present VoiceQualityVC, a tool that can manipulate the voice quality of both natural and synthesized speech using voice quality features including CPPS, H1-H2, and H1-A3. VoiceQualityVC is a research tool for perceptual experiments into voice quality and UX experiments for voice design. We perform an objective evaluation demonstrating the control of these features as well as subjective listening tests of the paralinguistic attributes of intimacy, valence, and investment. In these listening tests breathy voice was rated as more intimate and more invested than modal voice and creaky voice was rated as less intimate and less positive. The code and models can be found at https://github.com/Hfkml/VQVC.

Place, publisher, year, edition, pages
International Speech Communication Association, 2025
Keywords
Paralinguistics, Pragmatics, Voice conversion, Voice quality
National Category
Comparative Language Studies and Linguistics Natural Language Processing
Identifiers
urn:nbn:se:kth:diva-372784 (URN)10.21437/Interspeech.2025-902 (DOI)001585350500467 ()2-s2.0-105020036268 (Scopus ID)
Conference
26th Interspeech Conference 2025, Rotterdam, Netherlands, Kingdom of the, August 17-21, 2025
Note

QC 20251120

Available from: 2025-11-20 Created: 2025-11-20 Last updated: 2026-05-29Bibliographically approved
Marcinek, L., Beskow, J. & Gustafsson, J. (2024). A dual-control dialogue framework for human-robot interaction data collection: integrating human emotional and contextual awareness with conversational AI. In: International Conference of Social Robotics (ICSR 2024): . Paper presented at International Conference of Social Robotics (ICSR 2024), Odense, Denmark, 24-26 October, 2024.
Open this publication in new window or tab >>A dual-control dialogue framework for human-robot interaction data collection: integrating human emotional and contextual awareness with conversational AI
2024 (English)In: International Conference of Social Robotics (ICSR 2024), 2024Conference paper, Poster (with or without abstract) (Refereed)
Abstract [en]

This paper presents a dialogue framework designed to capture human-robot interactions enriched with human-level situational awareness. The system integrates advanced large language models with realtime human-in-the-loop control. Central to this framework is an interaction manager that oversees information flow, turn-taking, and prosody control of a social robot’s responses. A key innovation is the control interface, enabling a human operator to perform tasks such as emotion recognition and action detection through a live video feed. The operator also manages high-level tasks, like topic shifts or behaviour instructions.

Input from the operator is incorporated into the dialogue context managed by GPT-4o, thereby influencing the ongoing interaction. This allows for the collection of interactional data from an automated system that leverages human-level emotional and situational awareness. The audiovisual data will be used to explore the impact of situational awareness on user behaviors in task-oriented human-robot interaction.

National Category
Natural Language Processing
Research subject
Speech and Music Communication
Identifiers
urn:nbn:se:kth:diva-375300 (URN)
Conference
International Conference of Social Robotics (ICSR 2024), Odense, Denmark, 24-26 October, 2024
Note

QC 20260112

Available from: 2026-01-12 Created: 2026-01-12 Last updated: 2026-01-12Bibliographically approved
Francis, J., Székely, É. & Gustafsson, J. (2024). ConnecTone: A Modular AAC System Prototype with Contextual Generative Text Prediction and Style-Adaptive Conversational TTS. In: Interspeech 2024: . Paper presented at 25th Interspeech Conferece 2024, Kos Island, Greece, September 1-5, 2024 (pp. 1001-1002). International Speech Communication Association
Open this publication in new window or tab >>ConnecTone: A Modular AAC System Prototype with Contextual Generative Text Prediction and Style-Adaptive Conversational TTS
2024 (English)In: Interspeech 2024, International Speech Communication Association , 2024, p. 1001-1002Conference paper, Published paper (Refereed)
Abstract [en]

Recent developments in generative language modeling and conversational Text-to-Speech present transformative potential for enhancing Augmentative and Alternative Communication (AAC) devices. Practical application of these technologies requires extensive research and testing. To address this, we introduce ConnecTone, a modular platform designed for rapid integration and testing of language generation and speech technology. ConnecTone implements context-sensitive generative text prediction, using conversational context from Automatic Speech Recognition inputs. The system incorporates a neural TTS that supports interpolation between reading and spontaneous conversational styles, along with adjustable prosodic features. These speech characteristics are predicted using Large Language Models, but can be adjusted by users for individual needs. We anticipate ConnecTone will enable us to rapidly evaluate and implement innovations, thereby contributing to faster benefit delivery to AAC users.

Place, publisher, year, edition, pages
International Speech Communication Association, 2024
Keywords
AAC, Human-Computer Interaction, Speech Synthesis, TTS
National Category
Natural Language Processing Computer Sciences Human Computer Interaction
Identifiers
urn:nbn:se:kth:diva-358873 (URN)2-s2.0-85214814511 (Scopus ID)
Conference
25th Interspeech Conferece 2024, Kos Island, Greece, September 1-5, 2024
Note

QC 20250124

Available from: 2025-01-23 Created: 2025-01-23 Last updated: 2025-01-24Bibliographically approved
Wang, S., Székely, É. & Gustafsson, J. (2024). Contextual Interactive Evaluation of TTS Models in Dialogue Systems. In: Interspeech 2024: . Paper presented at 25th Interspeech Conferece 2024, Kos Island, Greece, Sep 1 2024 - Sep 5 2024 (pp. 2965-2969). International Speech Communication Association
Open this publication in new window or tab >>Contextual Interactive Evaluation of TTS Models in Dialogue Systems
2024 (English)In: Interspeech 2024, International Speech Communication Association , 2024, p. 2965-2969Conference paper, Published paper (Refereed)
Abstract [en]

Evaluation of text-to-speech (TTS) models is currently dominated by Mean-Opinion-Score (MOS) listening test, but MOS has been increasingly questioned for its validity. MOS tests place listeners in a passive setup, in which they do not actively interact with the TTS and usually evaluate isolated utterances without context. Thus it gives no indication for how well a TTS model suits an interactive application like spoken dialogue system, in which the capability of generating appropriate speech in the dialogue context is paramount. We aim to take a first step towards addressing this shortcoming by evaluating several state-of-the-art neural TTS models, including one that adapts to dialogue context, in a custom-built spoken dialogue system. We present system design, experiment setup, and results. Our work is the first to evaluate TTS in contextual dialogue system interactions. We also discuss the shortcomings and future opportunities of the proposed evaluation paradigm.

Place, publisher, year, edition, pages
International Speech Communication Association, 2024
Keywords
evaluation methodology, human-computer interaction, spoken dialogue system, text-to-speech
National Category
Natural Language Processing Other Engineering and Technologies
Identifiers
urn:nbn:se:kth:diva-358876 (URN)10.21437/Interspeech.2024-1008 (DOI)001331850103017 ()2-s2.0-85214809755 (Scopus ID)
Conference
25th Interspeech Conferece 2024, Kos Island, Greece, Sep 1 2024 - Sep 5 2024
Note

QC 20250128

Available from: 2025-01-23 Created: 2025-01-23 Last updated: 2025-12-05Bibliographically approved
Organisations
Identifiers
ORCID iD: ORCID iD iconorcid.org/0000-0002-0397-6442

Search in DiVA

Show all publications