kth.sePublications KTH
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Improved Dysarthric Speech to Text Conversion via TTS Personalization
Hungarian Research Centre for Linguistics, HUN-REN, Hungary; Department of Telecommunications and Artificial Intelligence, Budapest University of Technology, Hungary.
KTH, School of Electrical Engineering and Computer Science (EECS), Intelligent systems, Speech, Music and Hearing, TMH.ORCID iD: 0000-0003-1175-840X
Department of Telecommunications and Artificial Intelligence, Budapest University of Technology, Hungary.
SpeechTex Ltd., Hungary; Hungarian Research Centre for Linguistics, HUN-REN, Hungary.
Show others and affiliations
2025 (English)In: 2025 33rd European Signal Processing Conference, EUSIPCO 2025 - Proceedings, Institute of Electrical and Electronics Engineers (IEEE) , 2025, p. 521-525Conference paper, Published paper (Refereed)
Abstract [en]

We present a case study on developing a customized speech-to-text system for a Hungarian speaker with severe dysarthria. State-of-the-art automatic speech recognition (ASR) models struggle with zero-shot transcription of dysarthric speech, yielding high error rates. To improve performance with limited real dysarthric data, we fine-tune an ASR model using synthetic speech generated via a personalized text-to-speech (TTS) system. We introduce a method for generating synthetic dysarthric speech with controlled severity by leveraging premorbidity recordings of the given speaker and speaker embedding interpolation, enabling ASR fine-tuning on a continuum of impairments. Fine-tuning on both real and synthetic dysarthric speech reduces the character error rate (CER) from 36-51% (zero-shot) to 7.3%. Our monolingual FastConformer_Hu ASR model significantly outperforms Whisper-turbo when fine-tuned on the same data, and the inclusion of synthetic speech contributes to an 18% relative CER reduction. These results highlight the potential of personalized ASR systems for improving accessibility for individuals with severe speech impairments.

Place, publisher, year, edition, pages
Institute of Electrical and Electronics Engineers (IEEE) , 2025. p. 521-525
Keywords [en]
automatic speech recognition, dysarthric speech, few-shot learning, text to speech synthesis
National Category
Natural Language Processing
Identifiers
URN: urn:nbn:se:kth:diva-385592DOI: 10.23919/EUSIPCO63237.2025.11226704Scopus ID: 2-s2.0-105029819672OAI: oai:DiVA.org:kth-385592DiVA, id: diva2:2086794
Conference
33rd European Signal Processing Conference, EUSIPCO 2025, Palermo, Italy, September 8-12, 2025
Note

Part of ISBN 9789464593624

QC 20260716

Available from: 2026-07-16 Created: 2026-07-16 Last updated: 2026-07-16Bibliographically approved

Open Access in DiVA

No full text in DiVA

Other links

Publisher's full textScopus

Authority records

Székely, Éva

Search in DiVA

By author/editor
Székely, Éva
By organisation
Speech, Music and Hearing, TMH
Natural Language Processing

Search outside of DiVA

GoogleGoogle Scholar

doi
urn-nbn

Altmetric score

doi
urn-nbn
Total: 14 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf