kth.sePublications KTH
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
On the Use of Self-Supervised Speech Representations in Spontaneous Speech Synthesis
KTH, School of Electrical Engineering and Computer Science (EECS), Speech, Music and Hearing.
KTH, School of Electrical Engineering and Computer Science (EECS), Speech, Music and Hearing.ORCID iD: 0000-0002-1643-1054
KTH, School of Electrical Engineering and Computer Science (EECS), Speech, Music and Hearing.ORCID iD: 0000-0002-0397-6442
KTH, School of Electrical Engineering and Computer Science (EECS), Speech, Music and Hearing.ORCID iD: 0000-0003-1175-840X
2023 (English)In: Proceedings of the 12th ISCA Speech Synthesis Workshop (SSW2023), Grenoble, France: International Speech Communication Association , 2023, p. 163-169Conference paper, Published paper (Refereed)
Abstract [en]

Self-supervised learning (SSL) speech representations learned from large amounts of diverse, mixed-quality speech data without transcriptions are gaining ground in many speech technology applications. Prior work has shown that SSL is an effective intermediate representation in two-stage text-to-speech (TTS) for both read and spontaneous speech. However, it is still not clear which SSL and which layer from each SSL model is most suited for spontaneous TTS. We address this shortcoming by extending the scope of comparison for SSL in spontaneous TTS to 6 different SSLs and 3 layers within each SSL. Furthermore, SSL has also shown potential in predicting the mean opinion scores (MOS) of synthesized speech, but this has only been done in read-speech MOS prediction. We extend an SSL-based MOS prediction framework previously developed for scoring read speech synthesis and evaluate its performance on synthesized spontaneous speech. All experiments are conducted twice on two different spontaneous corpora in order to find generalizable trends. Overall, we present comprehensive experimental results on the use of SSL in spontaneous TTS and MOS prediction to further quantify and understand how SSL can be used in spontaneous TTS. Audios samples: https://www.speech.kth.se/tts-demos/sp_ssl_tts.

Place, publisher, year, edition, pages
Grenoble, France: International Speech Communication Association , 2023. p. 163-169
Keywords [en]
spontaneous speech synthesis, text-to-speech, self-supervised learning, mean-opinion-score prediction
National Category
Computer Sciences Natural Language Processing
Research subject
Speech and Music Communication
Identifiers
URN: urn:nbn:se:kth:diva-386105DOI: 10.21437/SSW.2023-26OAI: oai:DiVA.org:kth-386105DiVA, id: diva2:2088163
Conference
12th ISCA Speech Synthesis Workshop, SSW 2023, Grenoble, France, August 26-28, 2023
Funder
Swedish Research Council, VR-2019-05003Swedish Research Council, VR-2020-02396Wallenberg AI, Autonomous Systems and Software Program (WASP)
Note

WASP_publications

QC 20260727

Available from: 2026-07-24 Created: 2026-07-24 Last updated: 2026-07-27Bibliographically approved

Open Access in DiVA

No full text in DiVA

Other links

Publisher's full text

Authority records

Wang, SiyangHenter, Gustav EjeGustafson, JoakimSzékely, Éva

Search in DiVA

By author/editor
Wang, SiyangHenter, Gustav EjeGustafson, JoakimSzékely, Éva
By organisation
Speech, Music and Hearing
Computer SciencesNatural Language Processing

Search outside of DiVA

GoogleGoogle Scholar

doi
urn-nbn

Altmetric score

doi
urn-nbn
Total: 8 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf