kth.sePublications KTH
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Yeah, Un, Oh: Continuous and Real-time Backchannel Prediction with Fine-tuning of Voice Activity Projection
Kyoto Univ, Grad Sch Informat, Kyoto, Japan.
Kyoto Univ, Grad Sch Informat, Kyoto, Japan.
KTH, School of Electrical Engineering and Computer Science (EECS), Speech, Music and Hearing.ORCID iD: 0000-0002-8579-1790
Kyoto Univ, Grad Sch Informat, Kyoto, Japan.
2025 (English)In: Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers): HUMAN LANGUAGE TECHNOLOGIES, VOL 1: LONG PAPERS / [ed] Chiruzzo, L Wang, L, Association for Computational Linguistics (ACL) , 2025, p. 7171-7181Conference paper, Published paper (Refereed)
Abstract [en]

In human conversations, short backchannel utterances such as "yeah" and "oh" play a crucial role in facilitating smooth and engaging dialogue. These backchannels signal attentiveness and understanding without interrupting the speaker, making their accurate prediction essential for creating more natural conversational agents. This paper proposes a novel method for real-time, continuous backchannel prediction using a fine-tuned Voice Activity Projection (VAP) model. While existing approaches have relied on turn-based or artificially balanced datasets, our approach predicts both the timing and type of backchannels in a continuous and frame-wise manner on unbalanced, real-world datasets. We first pre-train the VAP model on a general dialogue corpus to capture conversational dynamics and then fine-tune it on a specialized dataset focused on backchannel behavior. Experimental results demonstrate that our model outperforms baseline methods in both timing and type prediction tasks, achieving robust performance in real-time environments. This research offers a promising step toward more responsive and human-like dialogue systems, with implications for interactive spoken dialogue applications such as virtual assistants and robots.

Place, publisher, year, edition, pages
Association for Computational Linguistics (ACL) , 2025. p. 7171-7181
National Category
Natural Language Processing
Identifiers
URN: urn:nbn:se:kth:diva-378052DOI: 10.18653/v1/2025.naacl-long.367ISI: 001611654000048Scopus ID: 2-s2.0-105027382814OAI: oai:DiVA.org:kth-378052DiVA, id: diva2:2046860
Conference
2025 Conference of the North American Chapter of the Association for Computational Linguistics-NAACL, APR 29-MAY 04, 2025, MEXICO
Note

Part of ISBN 979-8-89176-189-6

QC 20260318

Available from: 2026-03-18 Created: 2026-03-18 Last updated: 2026-03-18Bibliographically approved

Open Access in DiVA

No full text in DiVA

Other links

Publisher's full textScopus

Authority records

Skantze, Gabriel

Search in DiVA

By author/editor
Skantze, Gabriel
By organisation
Speech, Music and Hearing
Natural Language Processing

Search outside of DiVA

GoogleGoogle Scholar

doi
urn-nbn

Altmetric score

doi
urn-nbn
Total: 31 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf