kth.sePublications KTH
Change search
Link to record
Permanent link

Direct link
Jiang, Bing'er
Alternative names
Publications (4 of 4) Show all publications
Inoue, K., Jiang, B., Ekstedt, E., Kawahara, T. & Skantze, G. (2024). Multilingual Turn-taking Prediction Using Voice Activity Projection. In: 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation, LREC-COLING 2024 - Main Conference Proceedings: . Paper presented at Joint 30th International Conference on Computational Linguistics and 14th International Conference on Language Resources and Evaluation, LREC-COLING 2024, Hybrid, May 20-25, 2024, Torino, Italy (pp. 11873-11883). European Language Resources Association (ELRA)
Open this publication in new window or tab >>Multilingual Turn-taking Prediction Using Voice Activity Projection
Show others...
2024 (English)In: 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation, LREC-COLING 2024 - Main Conference Proceedings, European Language Resources Association (ELRA) , 2024, p. 11873-11883Conference paper, Published paper (Refereed)
Abstract [en]

This paper investigates the application of voice activity projection (VAP), a predictive turn-taking model for spoken dialogue, on multilingual data, encompassing English, Mandarin, and Japanese. The VAP model continuously predicts the upcoming voice activities of participants in dyadic dialogue, leveraging a cross-attention Transformer to capture the dynamic interplay between participants. The results show that a monolingual VAP model trained on one language does not make good predictions when applied to other languages. However, a multilingual model, trained on all three languages, demonstrates predictive performance on par with monolingual models across all languages. Further analyses show that the multilingual model has learned to discern the language of the input signal. We also analyze the sensitivity to pitch, a prosodic cue that is thought to be important for turn-taking. Finally, we compare two different audio encoders, contrastive predictive coding (CPC) pre-trained on English, with a recent model based on multilingual wav2vec 2.0 (MMS).

Place, publisher, year, edition, pages
European Language Resources Association (ELRA), 2024
Keywords
Multilingual, Spoken Dialogue System, Turn-taking, Voice Activity Projection
National Category
Natural Language Processing General Language Studies and Linguistics Computer Sciences
Identifiers
urn:nbn:se:kth:diva-348790 (URN)2-s2.0-85195914079 (Scopus ID)
Conference
Joint 30th International Conference on Computational Linguistics and 14th International Conference on Language Resources and Evaluation, LREC-COLING 2024, Hybrid, May 20-25, 2024, Torino, Italy
Projects
tmh_turntaking
Note

Part of ISBN 978-249381410-4

QC 20241028

Available from: 2024-06-27 Created: 2024-06-27 Last updated: 2025-02-01Bibliographically approved
Inoue, K., Jiang, B., Ekstedt, E., Kawahara, T. & Skantze, G. (2024). Real-time and Continuous Turn-taking Prediction Using Voice Activity Projection. In: : . Paper presented at The 14th International Workshop on Spoken Dialogue Systems Technology (IWSDS), Sapporo, Japan, March 4-6, 2024.
Open this publication in new window or tab >>Real-time and Continuous Turn-taking Prediction Using Voice Activity Projection
Show others...
2024 (English)Conference paper, Oral presentation with published abstract (Refereed)
Abstract [en]

A demonstration of a real-time and continuous turn-taking prediction system is presented. The system is based on a voice activity projection (VAP) model, which directly maps dialogue stereo audio to future voice activities. The VAP model includes contrastive predictive coding (CPC) and self-attention transformers, followed by a cross-attention transformer. We examine the effect of the input context audio length and demonstrate that the proposed system can operate in real-time with CPU settings, with minimal performance degradation.

National Category
Natural Language Processing
Identifiers
urn:nbn:se:kth:diva-359141 (URN)10.48550/arXiv.2401.04868 (DOI)
Conference
The 14th International Workshop on Spoken Dialogue Systems Technology (IWSDS), Sapporo, Japan, March 4-6, 2024
Projects
tmh_turntaking
Note

QC 20250325

Available from: 2025-01-27 Created: 2025-01-27 Last updated: 2025-03-25Bibliographically approved
Jiang, B., Ekstedt, E. & Skantze, G. (2023). Response-conditioned Turn-taking Prediction. In: Findings of the Association for Computational Linguistics, ACL 2023: . Paper presented at 61st Annual Meeting of the Association for Computational Linguistics, ACL 2023, July 9-14, 2023, Toronto, Canada (pp. 12241-12248). Association for Computational Linguistics (ACL)
Open this publication in new window or tab >>Response-conditioned Turn-taking Prediction
2023 (English)In: Findings of the Association for Computational Linguistics, ACL 2023, Association for Computational Linguistics (ACL) , 2023, p. 12241-12248Conference paper, Published paper (Refereed)
Abstract [en]

Previous approaches to turn-taking and response generation in conversational systems have treated it as a two-stage process: First, the end of a turn is detected (based on conversation history), then the system generates an appropriate response. Humans, however, do not take the turn just because it is likely, but also consider whether what they want to say fits the position. In this paper, we present a model (an extension of TurnGPT) that conditions the end-of-turn prediction on both conversation history and what the next speaker wants to say. We find that our model consistently outperforms the baseline model on a variety of metrics. The improvement is most prominent in two scenarios where turn predictions can be ambiguous solely from the conversation history: 1) when the current utterance contains a statement followed by a question; 2) when the end of the current utterance semantically matches the response. Treating the turn-prediction and response-ranking as a one-stage process, our findings suggest that our model can be used as an incremental response ranker, which can be applied in various settings.

Place, publisher, year, edition, pages
Association for Computational Linguistics (ACL), 2023
National Category
Natural Language Processing
Identifiers
urn:nbn:se:kth:diva-350243 (URN)10.18653/v1/2023.findings-acl.776 (DOI)2-s2.0-85175451617 (Scopus ID)
Conference
61st Annual Meeting of the Association for Computational Linguistics, ACL 2023, July 9-14, 2023, Toronto, Canada
Projects
tmh_turntaking
Note

Part of ISBN 9781959429623

QC 20241028

Available from: 2024-07-11 Created: 2024-07-11 Last updated: 2025-02-07Bibliographically approved
Jiang, B., Ekstedt, E. & Skantze, G. (2023). What makes a good pause? Investigating the turn-holding effects of fillers. In: Proceedings 20th International Congress of Phonetic Sciences (ICPhS): . Paper presented at 20th International Congress of Phonetic Sciences (ICPhS). August 7-11 2023, Prague, Czech Republic (pp. 3512-3516). Prague: International Phonetic Association, Article ID 828.
Open this publication in new window or tab >>What makes a good pause? Investigating the turn-holding effects of fillers
2023 (English)In: Proceedings 20th International Congress of Phonetic Sciences (ICPhS), Prague: International Phonetic Association , 2023, p. 3512-3516, article id 828Conference paper, Published paper (Refereed)
Abstract [en]

Filled pauses (or fillers), such as uh and um, are frequent in spontaneous speech and can serve as a turn-holding cue for the listener, indicating that the current speaker is not done yet. In this paper, we use the recently proposed Voice Activity Projection (VAP) model, which is a deep learning model trained to predict the dynamics of conversation, to analyse the effects of filled pauses on the expected turn-hold probability. The results show that, while filled pauses do indeed have a turn-holding effect, it is perhaps not as strong as could be expected, probably due to the redundancy of other cues. We also find that the prosodic properties and position of the filler has a significant effect on the turn-hold probability. However, contrary to what has been suggested in previous work, there is no difference between uh and um in this regard.

Place, publisher, year, edition, pages
Prague: International Phonetic Association, 2023
Series
ICPhS Proceedings, ISSN 2412-0669
Keywords
Hesitation, fillers, turn-taking, spoken dialog, computational modelling
National Category
Natural Language Processing
Identifiers
urn:nbn:se:kth:diva-341383 (URN)
Conference
20th International Congress of Phonetic Sciences (ICPhS). August 7-11 2023, Prague, Czech Republic
Projects
tmh_turntaking
Note

Part of ISBN 978-80-908 114-2-3

QC 20241028

Available from: 2023-12-19 Created: 2023-12-19 Last updated: 2025-02-07Bibliographically approved
Organisations

Search in DiVA

Show all publications