kth.sePublications KTH
Change search
Link to record
Permanent link

Direct link
Alternative names
Publications (10 of 15) Show all publications
Moëll, B. & Sand Aronsson, F. (2026). High-accuracy prediction of mental health scores from English BERT embeddings trained on LLM-generated synthetic self-reports: a synthetic-only method development study. Frontiers in Digital Health, 7, Article ID 1694464.
Open this publication in new window or tab >>High-accuracy prediction of mental health scores from English BERT embeddings trained on LLM-generated synthetic self-reports: a synthetic-only method development study
2026 (English)In: Frontiers in Digital Health, E-ISSN 2673-253X, Vol. 7, article id 1694464Article in journal (Refereed) Published
Abstract [en]

Objective: To assess whether synthetic-only first-person clinical self-reports generated by a large language model (LLM) can support accurate prediction of standardized mental-health scores, enabling a privacy-preserving path for method development and rapid prototyping when real clinical text is unavailable. Methods: We prompted an LLM (Gemini 2.5; July 2025 snapshot) to produce English-language first-person narratives that are paired with target scores for three instruments—PHQ-9 (including suicidal ideation), LSAS, and PCL-5. No real patients or clinical notes were used. Narratives and labels were created synthetically and manually screened for coherence and label alignment. Each narrative was embedded using bert-base-uncased (mean-pooled 768-d vectors). We trained linear/regularized linear (Linear, Ridge, Lasso) and ensemble models (Random Forest, Gradient Boosting) for regression, and Logistic Regression/Random Forest for suicidal-ideation classification. Evaluation used 5-fold cross-validation (PHQ-9/SI) and 80/20 held-out splits (LSAS/PCL-5). Metrics: MSE, (Formula presented.), MAE; classification metrics are reported for SI. Results: Within the synthetic distribution, models fit the label–text signal strongly (e.g., PHQ-9 Ridge: MSE (Formula presented.), (Formula presented.) ; LSAS Gradient Boosting test: MSE (Formula presented.), (Formula presented.) ; PCL-5 Ridge test: MSE (Formula presented.), (Formula presented.)). Conclusions: LLM-generated self-reports encode a score-aligned signal that standard ML models can learn, indicating utility for privacy-preserving, synthetic-only prototyping. This is not a clinical tool: results do not imply generalization to real patient text. We clarify terminology (synthetic text vs. real text) and provide a roadmap for external validation, bias/fidelity assessment, and scope-limited deployment considerations before any clinical use.

Place, publisher, year, edition, pages
Frontiers Media SA, 2026
Keywords
BERT, digital mental health, large language models, LSAS, natural language processing, PCL-5, PHQ-9, privacy-preserving evaluation
National Category
Natural Language Processing
Identifiers
urn:nbn:se:kth:diva-376520 (URN)10.3389/fdgth.2025.1694464 (DOI)001667048700001 ()41586203 (PubMedID)2-s2.0-105028571312 (Scopus ID)
Note

QC 20260209

Available from: 2026-02-09 Created: 2026-02-09 Last updated: 2026-02-09Bibliographically approved
Moell, B. & Sand Aronsson, F. (2025). Automatic Evaluation of the Pataka Test Using Machine Learning and Audio Signal Processing. Acta Logopaedica, 2
Open this publication in new window or tab >>Automatic Evaluation of the Pataka Test Using Machine Learning and Audio Signal Processing
2025 (English)In: Acta Logopaedica, E-ISSN 2004-9048, Vol. 2Article in journal (Refereed) Published
Abstract [en]

This study presents an automated deep learning approach to evaluate the oral diadochokinesis, a widely used clinical tool for assessing syllable repetition speed in motor speech disorders. Addressing the limitations of manual assessments—including subjectivity, time constraints, and inter-rater variability—we developed a system leveraging the Wav2Vec2 speech recognition model, combined with audio preprocessing (resampling, mono conversion, and normalisation) and temporal alignment techniques for syllable detection. In an initial assessment of the developed method, the system was evaluated on 16 recordings from two healthy speakers, analysed by three speech and language pathologists (SLPs) and compared to ground truth measurements. Results demonstrated superior accuracy of the machine learning system, with a mean squared error (MSE) of 0.07, compared to 1.18 for human raters. Statistical analysis (Wilcoxon signed-rank test: p = 0.98 for model vs. p = 0.00043 for SLPs) confirmed the model’s alignment with ground truth. While the system occasionally missed syllables (1–2 per recording), its precision in calculating syllables per second (SPS) and temporal consistency highlights its potential as a supplementary clinical tool. Key innovations include a user-friendly offline interface for data security and visualisations (Mel spectrograms, timing evenness, and distinctness metrics) to support clinical interpretation. The present study is subject to certain limitations. The study’s methodology is constrained by a small and homogeneous sample. Separately, the developed system’s performance is limited by unresolved challenges in the detection of subtle articulation errors. Future work will expand validation to diverse populations, including speakers with dysarthria, and refine human-in-the-loop integration to mitigate missed syllables. This study underscores the feasibility of combining deep learning with signal processing to enhance objectivity in speech assessments, offering a scalable solution to standardise the oral diadochokinesis test while preserving clinical expertise.

Place, publisher, year, edition, pages
CLINTEC/Logopedi, Karolinska Institutet, 2025
National Category
Computer Sciences
Identifiers
urn:nbn:se:kth:diva-371730 (URN)10.58986/al.2025.41035 (DOI)
Note

QC 20251019

Available from: 2025-10-17 Created: 2025-10-17 Last updated: 2025-10-19Bibliographically approved
Moëll, B. (2025). Evaluation of Artificial Intelligence in the Medical Domain: Speech, Language and Applications. (Doctoral dissertation). Stockholm: KTH Royal Institute of Technology
Open this publication in new window or tab >>Evaluation of Artificial Intelligence in the Medical Domain: Speech, Language and Applications
2025 (English)Doctoral thesis, comprehensive summary (Other academic)
Abstract [en]

This doctoral thesis investigates the potential of advanced speech and languagetechnologies, driven by deep learning, to improve clinical diagnostics and patientcare, primarily within the Swedish healthcare context. The research encompasseseight key papers, which are presented across three main sections:(1) Data Capture and Machine Learning for Speech: This section explores the use ofmultimodal data and advanced speech processing techniques for clinical applications.It includes research on utilizing multimodal data capture (speech, gaze, and digitalpen input) from clinical interviews to identify potential digital biomarkers for theearly detection and differentiation of dementia (Paper A). It also develops anautomated deep learning system to evaluate the oral diadochokinesis test for motorspeech disorders, which demonstrates higher accuracy than human raters andproposes a human-in-the-loop clinical interface (Paper B). Furthermore, this sectionevaluates the performance of Automatic Speech Recognition (ASR) systems,comparing word error rates between native (L1) and non-native (L2) Swedishspeakers (Paper C), and investigates data augmentation techniques to improve ASRaccuracy for individuals with aphasia, demonstrating a path towards more inclusivetechnology (Paper D).(2) Evaluation of LLMs in the Medical Domain: This section focuses on establishingrobust methods for assessing Large Language Models (LLMs) within a medicalcontext. It details the development of a specialized Swedish Medical LLM Benchmark,comprising over 2600 questions across various medical domains, designed to assessLLM performance in a clinically relevant, language-specific manner (Paper E).Additionally, the medical reasoning capabilities of LLMs, such as DeepSeek R1, arerigorously assessed, focusing on their capacity for general medical diagnosticreasoning (Paper F).(3) Application and Best Practice for Working with AI in Healthcare: This sectionaddresses the practical, ethical, and user experience (UX) considerations forvimplementing AI in healthcare. It proposes a novel user interface paradigm throughan AI-powered journaling application designed for personal health management,illustrating a low-risk, user-centric approach to AI integration (Paper G).Complementing this, it develops harm reduction strategies for the thoughtful use ofLLMs in the medical domain, providing perspectives for both patients and cliniciansto maximize utility while mitigating risks, thereby establishing best practices forresponsible AI engagement (Paper H).Collectively, this work advances the field by providing new tools and methodologiesfor early disease detection using speech and multimodal data, establishing robustevaluation methods for ASR and LLMs in the medical domain, and offering pathwaysand frameworks for responsible, user-centered, and effective AI implementation inhealthcare.

Abstract [sv]

Denna doktorsavhandling undersöker potentialen hos avancerade tal- ochspråkteknologier, drivna av djupinlärning, för att förbättra klinisk diagnostik ochpatientvård, främst inom svensk hälso- och sjukvård. Forskningen omfattar åttacentrala artiklar, vilka presenteras inom tre huvudsakliga avsnitt:(1) Datainsamling och maskininlärning för tal: Detta avsnitt utforskar användningenav multimodal data och avancerade talbearbetningstekniker för kliniskatillämpningar. Det inkluderar forskning om användning av multimodaldatainsamling från kliniska intervjuer för att identifiera digitala biomarkörer fördemens (Artikel A). Vidare utvecklas ett automatiserat system med djupinlärning föratt utvärdera oral diadochokinesis-testet vid motoriska talrubbningar, vilket visarhögre noggrannhet än mänskliga bedömare och föreslår ett kliniskt gränssnitt medmänniska-i-loopen (Artikel B). Avsnittet utvärderar även prestandan hos system förautomatisk taligenkänning (ASR) genom att jämföra felkvoter mellan talare medsvenska som modersmål respektive andraspråk (Artikel C) och undersökerdataaugmenteringstekniker för att förbättra ASR-noggrannheten för personer medafasi (Artikel D).(2) Utvärdering av stora språkmodeller (LLM:er) inom det medicinska området:Detta avsnitt fokuserar på att etablera robusta metoder för att bedöma storaspråkmodeller (LLM:er) i en medicinsk kontext. Det beskriver utvecklingen av ettspecialiserat svenskt medicinskt LLM-benchmark, bestående av över 2600 frågorinom olika medicinska domäner, avsett att utvärdera LLM:ers prestanda på ettkliniskt relevant och språkspecifikt sätt (Artikel E). Därtill bedöms den medicinskaresonemangsförmågan hos LLM:er, såsom DeepSeek R1, noggrant, med fokus påderas kapacitet för generell medicinsk diagnostiskt resonerande (Artikel F).(3) Applikationer och bästa praxis för AI inom hälso- och sjukvård: Detta avsnittbehandlar praktiska, etiska och användarupplevelsemässiga (UX) överväganden vidimplementering av AI inom hälso- och sjukvården. Ett nyttviianvändargränssnittsparadigm föreslås genom en AI-driven applikation för att föra enpersonlig hälsodagbok. Den är utformad för personlig hälsohantering och illustreraren lågrisk, användarcentrerad strategi för AI-integration (Artikel G). Somkomplement utvecklas strategier för harm reduction för genomtänkt användning avLLM:er inom det medicinska området. Dessa strategier erbjuder perspektiv för bådepatienter och kliniker för att maximera nyttan och samtidigt minimera riskerna, ochetablerar därmed bästa praxis för ansvarsfullt AI-engagemang (Artikel H).Sammantaget bidrar detta arbete till forskningsfältet genom att tillhandahålla nyaverktyg och metoder för tidig sjukdomsdetektion med hjälp av tal- och multimodaldata, etablera robusta utvärderingsmetoder för ASR och LLM:er inom det medicinskaområdet, samt erbjuda vägledning och ramverk för en ansvarsfull, användarcentreradoch effektiv implementering av AI inom hälso- och sjukvården.

Place, publisher, year, edition, pages
Stockholm: KTH Royal Institute of Technology, 2025. p. xxi, 82
Series
TRITA-EECS-AVL ; 2025:83
Keywords
Large Language Models (LLMs), Automatic Speech Recognition (ASR), Neurodegenerative Disorders, Swedish Language, Clinical Diagnostics, AI Ethics, Medical Reasoning, Multimodal Data, Tal- och språkteknologi, maskininlärning, djupinlärning, automatisk taligenkänning (ASR), stora språkmodeller (LLM), medicinsk diagnostik, digitala biomarkörer, afasi, demens, hälso- och sjukvård, användarupplevelse (UX), harm reduction, AI-integration
National Category
Artificial Intelligence
Research subject
Speech and Music Communication
Identifiers
urn:nbn:se:kth:diva-371738 (URN)978-91-8106-404-9 (ISBN)
Public defence
2025-12-12, https://kth-se.zoom.us/j/69936124469, Kollegiesalen, Brinellvägen 8, Stockholm, 13:00 (English)
Opponent
Supervisors
Note

QC 20251022

Available from: 2025-10-22 Created: 2025-10-17 Last updated: 2025-11-13Bibliographically approved
Moëll, B. & Sand Aronsson, F. (2025). Harm Reduction Strategies for Thoughtful Use of Large Language Models in the Medical Domain: Perspectives for Patients and Clinicians. Journal of Medical Internet Research, 27, Article ID e75849.
Open this publication in new window or tab >>Harm Reduction Strategies for Thoughtful Use of Large Language Models in the Medical Domain: Perspectives for Patients and Clinicians
2025 (English)In: Journal of Medical Internet Research, E-ISSN 1438-8871, Vol. 27, article id e75849Article, review/survey (Refereed) Published
Abstract [en]

The integration of large language models (LLMs) into health care presents significant risks to patients and clinicians, inadequately addressed by current guidance. This paper adapts harm reduction principles from public health to medical LLMs, proposing a structured framework for mitigating these domain-specific risks while maximizing ethical utility. We outline tailored strategies for patients, emphasizing critical health literacy and output verification, and for clinicians, enforcing “human-in-the-loop” validation and bias-aware workflows. Key innovations include developing thoughtful use protocols that position LLMs as assistive tools requiring mandatory verification, establishing actionable institutional policies with risk-stratified deployment guidelines and patient disclaimers, and critically analyzing underaddressed regulatory, equity, and safety challenges. This research moves beyond theory to offer a practical roadmap, enabling stakeholders to ethically harness LLMs, balance innovation with accountability, and preserve core medical values: patient safety, equity, and trust in high-stakes health care settings.

Place, publisher, year, edition, pages
JMIR Publications Inc., 2025
Keywords
artificial intelligence, assistive technology, bias awareness, conversational AI, governance frameworks, health care innovation, human-in-the-loop, regulatory compliance, risk mitigation, trustworthiness, verification protocols
National Category
Medical Ethics Health Care Service and Management, Health Policy and Services and Health Economy
Identifiers
urn:nbn:se:kth:diva-368804 (URN)10.2196/75849 (DOI)001542093800002 ()40712151 (PubMedID)2-s2.0-105011835941 (Scopus ID)
Note

QC 20250821

Available from: 2025-08-21 Created: 2025-08-21 Last updated: 2025-10-17Bibliographically approved
Moëll, B. & Sand Aronsson, F. (2025). Journaling with large language models: a novel UX paradigm for AI-driven personal health management. Frontiers in Artificial Intelligence, 8, Article ID 1567580.
Open this publication in new window or tab >>Journaling with large language models: a novel UX paradigm for AI-driven personal health management
2025 (English)In: Frontiers in Artificial Intelligence, E-ISSN 2624-8212, Vol. 8, article id 1567580Article in journal (Refereed) Published
Abstract [en]

Introduction: The integration of large language models (LLMs) into personal health management presents transformative potential, but faces critical challenges in user experience (UX) design, ethical implementation, and clinical integration.

Method: This paper introduces a novel AI-driven journaling application, a functional prototype available open source, designed to encourage patient engagement through a natural language interface. This approach, termed “AI-assisted health journaling,” enables users to document health experiences in their own words while receiving real-time, context-aware feedback from an LLM. The prototype combines a personal health record with an LLM assistant, allowing for reflective self-monitoring and aiming to combine patient-generated data with clinical insights. Key innovations include a three-panel interface for seamless journaling, AI dialogue, and longitudinal tracking, alongside specialized modes for interacting with simulated healthcare expert personas.

Result: Preliminary insights from persona-based evaluations highlight the system's capacity to enhance health literacy through explainable AI responses while maintaining strict data localization and privacy controls. We propose five design principles for patient-centric AI health tools: (1) decoupling core functionality from LLM dependencies, (2) layered transparency in AI outputs, (3) adaptive consent for data sharing, (4) clinician-facing data summarization, and (5) compliance-first architecture.

Discussion: By transforming unstructured patient narratives into structured insights through natural language processing, this approach demonstrates how journaling interfaces could serve as a critical middleware layer in healthcare ecosystems-empowering patients as active partners in care while preserving clinical oversight. Future research directions emphasize the need for rigorous trials evaluating impacts on care continuity, patient-provider communication, and long-term health outcomes across diverse populations.

Place, publisher, year, edition, pages
Frontiers Media SA, 2025
Keywords
AI-driven journaling, data privacy, explainable AI, health literacy, large language models (LLMs), medical AI, natural language processing (NLP), patient engagement
National Category
Computer Sciences Human Computer Interaction
Identifiers
urn:nbn:se:kth:diva-369514 (URN)10.3389/frai.2025.1567580 (DOI)001523808200001 ()40630834 (PubMedID)2-s2.0-105010962742 (Scopus ID)
Note

QC 20250911

Available from: 2025-09-11 Created: 2025-09-11 Last updated: 2025-10-17Bibliographically approved
Moell, B., Aronsson, F. S. & Akbar, S. (2025). Medical reasoning in LLMs: an in-depth analysis of DeepSeek R1. Frontiers in Artificial Intelligence, 8, Article ID 1616145.
Open this publication in new window or tab >>Medical reasoning in LLMs: an in-depth analysis of DeepSeek R1
2025 (English)In: Frontiers in Artificial Intelligence, E-ISSN 2624-8212, Vol. 8, article id 1616145Article in journal (Refereed) Published
Abstract [en]

Introduction The integration of large language models (LLMs) into healthcare holds immense promise, but also raises critical challenges, particularly regarding the interpretability and reliability of their reasoning processes. While models like DeepSeek R1-which incorporates explicit reasoning steps-show promise in enhancing performance and explainability, their alignment with domain-specific expert reasoning remains understudied.Methods This paper evaluates the medical reasoning capabilities of DeepSeek R1, comparing its outputs to the reasoning patterns of medical domain experts.Results Through qualitative and quantitative analyses of 100 diverse clinical cases from the MedQA dataset, we demonstrate that DeepSeek R1 achieves 93% diagnostic accuracy and shows patterns of medical reasoning. Analysis of the seven error cases revealed several recurring errors: anchoring bias, difficulty integrating conflicting data, limited consideration of alternative diagnoses, overthinking, incomplete knowledge, and prioritizing definitive treatment over crucial intermediate steps.Discussion These findings highlight areas for improvement in LLM reasoning for medical applications. Notably the length of reasoning was important with longer responses having a higher probability for error. The marked disparity in reasoning length suggests that extended explanations may signal uncertainty or reflect attempts to rationalize incorrect conclusions. Shorter responses (e.g., under 5,000 characters) were strongly associated with accuracy, providing a practical threshold for assessing confidence in model-generated answers. Beyond observed reasoning errors, the LLM demonstrated sound clinical judgment by systematically evaluating patient information, forming a differential diagnosis, and selecting appropriate treatment based on established guidelines, drug efficacy, resistance patterns, and patient-specific factors. This ability to integrate complex information and apply clinical knowledge highlights the potential of LLMs for supporting medical decision-making through artificial medical reasoning.

Place, publisher, year, edition, pages
Frontiers Media SA, 2025
Keywords
LLM, medical reasoning, DeepSeek R1, AI in medicine, reasoning models, medical benchmarking
National Category
Computer Sciences
Identifiers
urn:nbn:se:kth:diva-371001 (URN)10.3389/frai.2025.1616145 (DOI)001521129800001 ()40607450 (PubMedID)2-s2.0-105009608977 (Scopus ID)
Note

QC 20251003

Available from: 2025-10-03 Created: 2025-10-03 Last updated: 2025-10-17Bibliographically approved
Moell, B., Farestam, F. & Beskow, J. (2025). Swedish Medical LLM Benchmark: Development and evaluation of a framework for assessing large language models in the Swedish medical domain. Frontiers in Artificial Intelligence, 8, Article ID 1557920.
Open this publication in new window or tab >>Swedish Medical LLM Benchmark: Development and evaluation of a framework for assessing large language models in the Swedish medical domain
2025 (English)In: Frontiers in Artificial Intelligence, E-ISSN 2624-8212, Vol. 8, article id 1557920Article in journal (Refereed) Published
Abstract [en]

Introduction: We present the Swedish Medical LLM Benchmark (SMLB), an evaluation framework for assessing large language models (LLMs) in the Swedish medical domain.

Method: The SMLB addresses the lack of language-specific, clinically relevant benchmarks by incorporating four datasets: translated PubMedQA questions, Swedish Medical Exams, Emergency Medicine scenarios, and General Medicine cases.

Result: Our evaluation of 18 state-of-the-art LLMs reveals GPT-4-turbo, Claude- 3.5 (October 2023), and the o3model as top performers, demonstrating a strong alignment between medical reasoning and general language understanding capabilities. Hybrid systems incorporating retrieval-augmented generation (RAG) improved accuracy for clinical knowledge questions, highlighting promising directions for safe implementation.

Discussion: The SMLB provides not only an evaluation tool but also reveals fundamental insights about LLM capabilities and limitations in Swedish healthcare applications, including significant performance variations between models. By open-sourcing the benchmark, we enable transparent assessment of medical LLMs while promoting responsible development through community-driven refinement. This study emphasizes the critical need for rigorous evaluation frameworks as LLMs become increasingly integrated into clinical workflows, particularly in non-English medical contexts where linguistic and cultural specificity are paramount.

 

Place, publisher, year, edition, pages
Frontiers Media SA, 2025
National Category
Natural Language Processing
Identifiers
urn:nbn:se:kth:diva-371731 (URN)10.3389/frai.2025.1557920 (DOI)001536176500001 ()40718621 (PubMedID)2-s2.0-105011480129 (Scopus ID)
Note

QC 20251019

Available from: 2025-10-17 Created: 2025-10-17 Last updated: 2025-11-13Bibliographically approved
Mehta, S., Deichler, A., O'Regan, J., Moëll, B., Beskow, J., Henter, G. E. & Alexanderson, S. (2024). Fake it to make it: Using synthetic data to remedy the data shortage in joint multi-modal speech-and-gesture synthesis. In: Proceedings - 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, CVPRW 2024: . Paper presented at 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, CVPRW 2024, Seattle, United States of America, Jun 16 2024 - Jun 22 2024 (pp. 1952-1964). Institute of Electrical and Electronics Engineers (IEEE)
Open this publication in new window or tab >>Fake it to make it: Using synthetic data to remedy the data shortage in joint multi-modal speech-and-gesture synthesis
Show others...
2024 (English)In: Proceedings - 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, CVPRW 2024, Institute of Electrical and Electronics Engineers (IEEE) , 2024, p. 1952-1964Conference paper, Published paper (Refereed)
Abstract [en]

Although humans engaged in face-to-face conversation simultaneously communicate both verbally and non-verbally, methods for joint and unified synthesis of speech audio and co-speech 3D gesture motion from text are a new and emerging field. These technologies hold great promise for more human-like, efficient, expressive, and robust synthetic communication, but are currently held back by the lack of suitably large datasets, as existing methods are trained on parallel data from all constituent modalities. Inspired by student-teacher methods, we propose a straightforward solution to the data shortage, by simply synthesising additional training material. Specifically, we use uni-modal synthesis models trained on large datasets to create multi-modal (but synthetic) parallel training data, and then pre-train a joint synthesis model on that material. In addition, we propose a new synthesis architecture that adds better and more controllable prosody modelling to the state-of-the-art method in the field. Our results confirm that pre-training on large amounts of synthetic data improves the quality of both the speech and the motion synthesised by the multi-modal model, with the proposed architecture yielding further benefits when pre-trained on the synthetic data.

Place, publisher, year, edition, pages
Institute of Electrical and Electronics Engineers (IEEE), 2024
Keywords
gesture synthesis, motion synthesis, multimodal synthesis, synthetic data, text-to-speech-and-motion, training-on-generated-data
National Category
Signal Processing Natural Language Processing
Identifiers
urn:nbn:se:kth:diva-367174 (URN)10.1109/CVPRW63382.2024.00201 (DOI)001327781702011 ()2-s2.0-85202828403 (Scopus ID)
Conference
2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, CVPRW 2024, Seattle, United States of America, Jun 16 2024 - Jun 22 2024
Note

Part of ISBN 9798350365474

QC 20250715

Available from: 2025-07-15 Created: 2025-07-15 Last updated: 2025-08-13Bibliographically approved
Moell, B., O'Regan, J., Mehta, S., Kirkland, A., Lameris, H., Gustafsson, J. & Beskow, J. (2022). Speech Data Augmentation for Improving Phoneme Transcriptions of Aphasic Speech Using Wav2Vec 2.0 for the PSST Challenge. In: Dimitrios Kokkinakis, Charalambos K. Themistocleous, Kristina Lundholm Fors, Athanasios Tsanas, Kathleen C. Fraser (Ed.), The RaPID4 Workshop: Resources and ProcessIng of linguistic, para-linguistic and extra-linguistic Data from people with various forms of cognitive/psychiatric/developmental impairments. Paper presented at 4th RaPID Workshop: Resources and Processing of Linguistic, Para-Linguistic and Extra-Linguistic Data from People with Various Forms of Cognitive/Psychiatric/Developmental Impairments, RAPID 2022, Marseille, France, Jun 25 2022 (pp. 62-70). Marseille, France
Open this publication in new window or tab >>Speech Data Augmentation for Improving Phoneme Transcriptions of Aphasic Speech Using Wav2Vec 2.0 for the PSST Challenge
Show others...
2022 (English)In: The RaPID4 Workshop: Resources and ProcessIng of linguistic, para-linguistic and extra-linguistic Data from people with various forms of cognitive/psychiatric/developmental impairments / [ed] Dimitrios Kokkinakis, Charalambos K. Themistocleous, Kristina Lundholm Fors, Athanasios Tsanas, Kathleen C. Fraser, Marseille, France, 2022, p. 62-70Conference paper, Published paper (Refereed)
Abstract [en]

As part of the PSST challenge, we explore how data augmentations, data sources, and model size affect phoneme transcription accuracy on speech produced by individuals with aphasia. We evaluate model performance in terms of feature error rate (FER) and phoneme error rate (PER). We find that data augmentations techniques, such as pitch shift, improve model performance. Additionally, increasing the size of the model decreases FER and PER. Our experiments also show that adding manually-transcribed speech from non-aphasic speakers (TIMIT) improves performance when Room Impulse Response is used to augment the data. The best performing model combines aphasic and non-aphasic data and has a 21.0% PER and a 9.2% FER, a relative improvement of 9.8% compared to the baseline model on the primary outcome measurement. We show that data augmentation, larger model size, and additional non-aphasic data sources can be helpful in improving automatic phoneme recognition models for people with aphasia.

Place, publisher, year, edition, pages
Marseille, France: , 2022
Keywords
aphasia, data augmentation, phoneme transcription, phonemes, speech, speech data augmentation, wav2vec 2.0
National Category
Other Electrical Engineering, Electronic Engineering, Information Engineering
Research subject
Speech and Music Communication
Identifiers
urn:nbn:se:kth:diva-314262 (URN)2-s2.0-85145876107 (Scopus ID)
Conference
4th RaPID Workshop: Resources and Processing of Linguistic, Para-Linguistic and Extra-Linguistic Data from People with Various Forms of Cognitive/Psychiatric/Developmental Impairments, RAPID 2022, Marseille, France, Jun 25 2022
Note

QC 20220815

Available from: 2022-06-17 Created: 2022-06-17 Last updated: 2025-10-17Bibliographically approved
Lameris, H., Mehta, S., Henter, G. E., Kirkland, A., Moëll, B., O'Regan, J., . . . Székely, É. (2022). Spontaneous Neural HMM TTS with Prosodic Feature Modification. In: Proceedings of Fonetik 2022: . Paper presented at Fonetik 2022, Stockholm 13-15 May, 202.
Open this publication in new window or tab >>Spontaneous Neural HMM TTS with Prosodic Feature Modification
Show others...
2022 (English)In: Proceedings of Fonetik 2022, 2022Conference paper, Published paper (Other academic)
Abstract [en]

Spontaneous speech synthesis is a complex enterprise, as the data has large variation, as well as speech disfluencies nor-mally omitted from read speech. These disfluencies perturb the attention mechanism present in most Text to Speech (TTS) sys-tems. Explicit modelling of prosodic features has enabled intu-itive prosody modification of synthesized speech. Most pros-ody-controlled TTS, however, has been trained on read-speech data that is not representative of spontaneous conversational prosody. The diversity in prosody in spontaneous speech data allows for more wide-ranging data-driven modelling of pro-sodic features. Additionally, prosody-controlled TTS requires extensive training data and GPU time which limits accessibil-ity. We use neural HMM TTS as it reduces the parameter size and can achieve fast convergence with stable alignments for spontaneous speech data. We modify neural HMM TTS to ena-ble prosodic control of the speech rate and fundamental fre-quency. We perform subjective evaluation of the generated speech of English and Swedish TTS models and objective eval-uation for English TTS. Subjective evaluation showed a signif-icant improvement in naturalness for Swedish for the mean prosody compared to a baseline with no prosody modification, and the objective evaluation showed greater variety in the mean of the per-utterance prosodic features.

National Category
Other Computer and Information Science Specific Languages
Identifiers
urn:nbn:se:kth:diva-313156 (URN)
Conference
Fonetik 2022, Stockholm 13-15 May, 202
Funder
Swedish Research Council, 2019-05003
Note

QC 20220726

Available from: 2022-05-31 Created: 2022-05-31 Last updated: 2024-03-15Bibliographically approved
Organisations
Identifiers
ORCID iD: ORCID iD iconorcid.org/0000-0002-6529-1211

Search in DiVA

Show all publications