kth.sePublikationer KTH
Ändra sökning
Länk till posten
Permanent länk

Direktlänk
Publikationer (10 of 114) Visa alla publikationer
Cai, H., Ternström, S., Chaffanjon, P. & Henrich Bernardoni, N. (2026). Effects on Voice Quality of Thyroidectomy: A Qualitative and Quantitative Study Using Voice Maps. Journal of Voice, 40(4), 1249.e21-1249.e42
Öppna denna publikation i ny flik eller fönster >>Effects on Voice Quality of Thyroidectomy: A Qualitative and Quantitative Study Using Voice Maps
2026 (Engelska)Ingår i: Journal of Voice, ISSN 0892-1997, E-ISSN 1873-4588, Vol. 40, nr 4, s. 1249.e21-1249.e42Artikel i tidskrift (Refereegranskat) Published
Abstract [en]

Objectives: This study aims to explore the effects of thyroidectomy—a surgical intervention involving the removal of the thyroid gland—on voice quality, as represented by acoustic and electroglottographic measures. Given the thyroid gland's proximity to the inferior and superior laryngeal nerves, thyroidectomy carries a potential risk of affecting vocal function. While earlier studies have documented effects on the voice range, few studies have looked at voice quality after thyroidectomy. Since voice quality effects could manifest in many ways, that a priori are unknown, we wish to apply an exploratory approach that collects many data points from several metrics.

Methods: A voice-mapping analysis paradigm was applied retrospectively on a corpus of spoken and sung sentences produced by patients who had thyroid surgery. Voice quality changes were assessed objectively for 57 patients prior to surgery and 2 months after surgery, by making comparative voice maps, pre- and post-intervention, of six acoustic and electroglottographic (EGG) metrics.

Results: After thyroidectomy, statistically significant changes consistent with a worsening of voice quality were observed in most metrics. For all individual metrics, however, the effect sizes were too small to be clinically relevant. Statistical clustering of the metrics helped to clarify the nature of these changes. While partial thyroidectomy demonstrated greater uniformity than did total thyroidectomy, the type of perioperative damage had no discernible impact on voice quality.ConclusionsChanges in voice quality after thyroidectomy were related mostly to increased phonatory instability in both the acoustic and EGG metrics. Clustered voice metrics exhibited a higher correlation to voice complaints than did individual voice metrics.

Ort, förlag, år, upplaga, sidor
Elsevier, 2026
Nyckelord
thyroidectomy, voice quality, electroglottography, voice classification, voice mapping
Nationell ämneskategori
Oto-rino-laryngologi
Forskningsämne
Tal- och musikkommunikation
Identifikatorer
urn:nbn:se:kth:diva-346224 (URN)10.1016/j.jvoice.2024.03.012 (DOI)38714436 (PubMedID)2-s2.0-85192255370 (Scopus ID)
Forskningsfinansiär
KTH, 6308
Anmärkning

QC 20260710

Tillgänglig från: 2024-05-07 Skapad: 2024-05-07 Senast uppdaterad: 2026-07-10Bibliografiskt granskad
Cai, H. & Ternström, S. (2025). A WaveNet-based model for predicting the electroglottographic signal from the acoustic voice signal. Journal of the Acoustical Society of America, 157(4), 3033-3044
Öppna denna publikation i ny flik eller fönster >>A WaveNet-based model for predicting the electroglottographic signal from the acoustic voice signal
2025 (Engelska)Ingår i: Journal of the Acoustical Society of America, ISSN 0001-4966, E-ISSN 1520-8524, Vol. 157, nr 4, s. 3033-3044Artikel i tidskrift (Refereegranskat) Published
Abstract [en]

The electroglottographic (EGG) signal offers a non-invasive approach to analyze phonation. It is known, if not obvious, that the onset of vocal fold contacting has a substantial effect on how the vocal folds vibrate and on the quality of the voice. Given that the presence or absence of vocal fold contacting has major consequences also for the interpretation of acoustic metrics, it is compelling to consider the possibility of predicting EGG signals directly from the microphone speech signal. This retrospective study presents a neural network model for EGG signal estimation utilizing a WaveNet architecture augmented with a self-attention mechanism. The model was trained on an existing dataset that comprehensively recorded participants' full voice range. The proposed model effectively captures the temporal dynamics and morphological characteristics of normophonic EGG waveforms, achieving outputs that closely resemble the ground truth in terms of EGG waveshape and extracted EGG metrics. For evaluation, voice mapping was used to display the distribution similarities of extracted metrics from predicted and ground truth EGG waveforms. The model exhibits proficiency in accurately estimating EGG signals in areas of stable and contacting voicing but displays reduced accuracy in transitional and breathy phonatory conditions.

Ort, förlag, år, upplaga, sidor
American Institute of Physics (AIP), 2025
Nyckelord
Phonetics, Vocalization, Vocal folds, Microphones, Speech analysis, Speech processing systems, Electroglottography, Acoustic signal processing, Artificial neural networks
Nationell ämneskategori
Oto-rino-laryngologi Medicinsk instrumentering Signalbehandling
Forskningsämne
Tal- och musikkommunikation
Identifikatorer
urn:nbn:se:kth:diva-362580 (URN)10.1121/10.0036514 (DOI)001472395500002 ()40249176 (PubMedID)2-s2.0-105003174138 (Scopus ID)
Anmärkning

A precursor to this article was included in Huanchen Cai's doctoral thesis. This is the revised and accepted version.

QC 20250425

Tillgänglig från: 2025-04-20 Skapad: 2025-04-20 Senast uppdaterad: 2025-12-08Bibliografiskt granskad
Ternström, S. & Pabon, P. (2025). From Voice Signals to Voice Maps. International Journal of Voice Sciences
Öppna denna publikation i ny flik eller fönster >>From Voice Signals to Voice Maps
2025 (Engelska)Ingår i: International Journal of Voice Sciences, E-ISSN 3054-4343Artikel i tidskrift (Refereegranskat) Epub ahead of print
Abstract [en]

This article is intended as an introductory tutorial for technically inclined clinicians, vocologists and voice pedagogues who want to understand the principles and potentials of voice mapping. Voice mapping has its origins in the Voice Range Profile, or phonetogram, but it is less concerned with the extremes of the voice range, and more with what happens within a relevant range of the voice. It is a voice instrumentation paradigm that is intended to improve the evidential value of voice measurements. It exposes and automatically accounts for the strong co-variation that most voice metrics exhibit with fundamental frequency and sound level. Very many data points are automatically collected in a short time, and their means are mapped by colour onto maps. This results in a robust representation of voice status and function. While individual voices are very different, a voice map’s appearance is reproducible within individuals. Comparing maps across interventions gives rich information, even on subtle changes in a voice. Further, by automatically clustering multiple metrics, phonation types can be identified and mapped automatically, which can increase the clinical relevance, and facilitate a better understanding of voice data in general.

Ort, förlag, år, upplaga, sidor
Hildesheim/Holzminden/Göttingen: Paradigm Publishers, 2025
Nyckelord
Voice map, voice range profile, voice measurement, electroglottography, clinical evidence
Nationell ämneskategori
Medicinsk instrumentering Oto-rino-laryngologi Signalbehandling
Forskningsämne
Tal- och musikkommunikation
Identifikatorer
urn:nbn:se:kth:diva-376855 (URN)10.2478/ijvs-2025-0002 (DOI)
Projekt
Språkbanken Tal; HumInfra
Anmärkning

QC 20260218

Tillgänglig från: 2026-02-18 Skapad: 2026-02-18 Senast uppdaterad: 2026-02-18Bibliografiskt granskad
Park, M., Ontakhrai, S., Kittimathaveenan, K., Alfredsson, J. & Ternström, S. (2025). How to make closed-back headphones transparent for avocalist’s own direct sound. In: : . Paper presented at AES 159th Convention 2025 October 23–25, Long Beach, CA, USA (pp. 8). Audio Engineering Society, Inc., Article ID 371.
Öppna denna publikation i ny flik eller fönster >>How to make closed-back headphones transparent for avocalist’s own direct sound
Visa övriga...
2025 (Engelska)Konferensbidrag, Publicerat paper (Refereegranskat)
Abstract [en]

In stage acoustics research, it is common to use virtual acoustic environments over headphones to simulate various room conditions for musicians. When experiments are conducted in suboptimal physical environments (e.g., withoutan anechoic chamber), it is often challenging to reduce the inherent reverberation of the test room while ensuring that musicians can hear their own direct sound through the headphones as if the headphones were transparent. In the present study, two methods were developed and tested using an acoustic dummy head, with the aim of faithfully replicating the direct sound of a solo singer over a pair of closed-back headphones. The results showed that the first method - creating and applying an exact finite-impulse-response (FIR) filter - may lead to undesirable effects, primarily due to the inherent delay of the playback system. The second method, which utilized a multiband equalizer, proved more effective when evaluated with stationary broadband noise. For non-stationary, real-world sounds such as singing and speaking voices, the comparison between the reference sound and the signal processed by the multiband equalizer remained reasonably accurate, with differences typically less than ~1 dB across much of the frequency range. Future work may further refine and evaluate the proposed methods through listening tests.

Ort, förlag, år, upplaga, sidor
Audio Engineering Society, Inc., 2025
Serie
AES E-Library
Nyckelord
singing, headphones, hearing-of-self, acoustic transparency
Nationell ämneskategori
Signalbehandling Musik Annan elektroteknik och elektronik
Forskningsämne
Tal- och musikkommunikation
Identifikatorer
urn:nbn:se:kth:diva-373098 (URN)
Konferens
AES 159th Convention 2025 October 23–25, Long Beach, CA, USA
Anmärkning

AES 159th Convention Express Paper 371.

This study was supported by the King Mongkut’s Institute of Technology Ladkrabang research grant (KREF046818), the Karl Engver Foundation, and OCSC, Thailand. 

QC 20251219

Tillgänglig från: 2025-11-18 Skapad: 2025-11-18 Senast uppdaterad: 2025-12-19Bibliografiskt granskad
Herbst, C. T., Tokuda, I. T., Nishimura, T., Ternström, S., Ossio, V., Levy, M., . . . Dunn, J. C. (2025). ‘Monkey yodels’—frequency jumps in New World monkey vocalizations greatly surpass human vocal register transitions. Philosophical Transactions of the Royal Society of London. Biological Sciences, 380(1923), Article ID 20240005.
Öppna denna publikation i ny flik eller fönster >>‘Monkey yodels’—frequency jumps in New World monkey vocalizations greatly surpass human vocal register transitions
Visa övriga...
2025 (Engelska)Ingår i: Philosophical Transactions of the Royal Society of London. Biological Sciences, ISSN 0962-8436, E-ISSN 1471-2970, Vol. 380, nr 1923, artikel-id 20240005Artikel i tidskrift (Refereegranskat) Published
Abstract [en]

We investigated the causal basis of abrupt frequency jumps in a unique database of New World monkey vocalizations. We used a combination of acoustic and electroglottographic recordings in vivo , excised larynx investigations of vocal fold dynamics, and computational modelling. We particularly attended to the contribution of the vocal membranes: thin upward extensions of the vocal folds found in most primates but absent in humans. In three of the six investigated species, we observed two distinct modes of vocal fold vibration. The first, involving vocal fold vibration alone, produced low-frequency oscillations, and is analogous to that underlying human phonation. The second, incorporating the vocal membranes, resulted in much higher-frequency oscillation. Abrupt fundamental frequency shifts were observed in all three datasets. While these data are reminiscent of the rapid transitions in frequency observed in certain human singing styles (e.g. yodelling), the frequency jumps are considerably larger in the nonhuman primates studied. Our data suggest that peripheral modifications of vocal anatomy provide an important source of variability and complexity in the vocal repertoires of nonhuman primates. We further propose that the call repertoire is crucially related to a species’ ability to vocalize with different laryngeal mechanisms, analogous to human vocal registers.

Ort, förlag, år, upplaga, sidor
The Royal Society, 2025
Nyckelord
vocal membrane, laryngeal mechanism, call repertoire, NLP vocalization, fundamental frequency contol
Nationell ämneskategori
Oto-rino-laryngologi Teknisk mekanik Strukturbiologi
Forskningsämne
Tal- och musikkommunikation
Identifikatorer
urn:nbn:se:kth:diva-219581 (URN)10.1098/rstb.2024.0005 (DOI)001461623200021 ()40176522 (PubMedID)2-s2.0-105001836522 (Scopus ID)
Anmärkning

QC 20250520

Tillgänglig från: 2025-04-03 Skapad: 2025-04-03 Senast uppdaterad: 2025-05-20Bibliografiskt granskad
Capobianco, S., Björck, G., Forli, F., Berrettini, S. & Ternström, S. (2025). Voice mapping in clinical practice: tracking objective changes after injection laryngoplasty. Otorinolaryngologie a foniatrie, 74(S1), s28-s28
Öppna denna publikation i ny flik eller fönster >>Voice mapping in clinical practice: tracking objective changes after injection laryngoplasty
Visa övriga...
2025 (Engelska)Ingår i: Otorinolaryngologie a foniatrie, ISSN 1210-7867, Vol. 74, nr S1, s. s28-s28Artikel i tidskrift (Refereegranskat) Published
Nationell ämneskategori
Oto-rino-laryngologi
Identifikatorer
urn:nbn:se:kth:diva-374602 (URN)10.48095/ccorl2025s1_46 (DOI)
Anmärkning

QC 20251219

Tillgänglig från: 2025-12-19 Skapad: 2025-12-19 Senast uppdaterad: 2025-12-19Bibliografiskt granskad
Capobianco, S., Björck, G., Forli, F., Bruschini, L., Nacci, A. & Ternström, S. (2025). Voice mapping in clinical practice: Tracking objective changes after injection laryngoplasty. In: Frassineti, L Lanata, A Manfredi, C (Ed.), Models and analysis of vocal emissions for biomedical applications: . Paper presented at 14th International Workshop on MODELS AND ANALYSIS OF VOCAL EMISSIONS FOR BIOMEDICAL APPLICATIONS-MAVEBA, DEC 16-17, 2025, Firenze, ITALY (pp. 15-18). Firenze Univ Press, 139
Öppna denna publikation i ny flik eller fönster >>Voice mapping in clinical practice: Tracking objective changes after injection laryngoplasty
Visa övriga...
2025 (Engelska)Ingår i: Models and analysis of vocal emissions for biomedical applications / [ed] Frassineti, L Lanata, A Manfredi, C, Firenze Univ Press , 2025, Vol. 139, s. 15-18Konferensbidrag, Publicerat paper (Refereegranskat)
Abstract [en]

Objective: to explore the use of voice mapping for assessing changes in phonatory function following injection laryngoplasty in patients with unilateral vocal fold paralysis (UVFP). Materials and methods: Two patient cohorts were analyzed. Cohort 1 (N=8) received in-office injections of hyaluronic acid or calcium hydroxylapatite, with voice recordings acquired before and immediately after treatment. Cohort 2 (N=4) underwent autologous fat injection under general anesthesia, with follow-ups at 1 and 3 months. All patients completed standard speech tasks with simultaneous acquisition of acoustic and electroglottographic (EGG) signals. Voice maps were computed using the FonaDyn system. Perceptual GRBAS ratings were provided by three blinded expert raters. Results: Voice mapping was feasible in all patients and revealed consistent treatment effects. Across both cohorts, the cycle-rate sample entropy (CSE) decreased, while the normalized peak dEGG (Q(Delta)) and the Index of Contacting (I-c) both increased, indicating improved phonatory stability and vocal fold contact. Perceptual ratings showed corresponding reductions in breathiness and overall dysphonia. Conclusions: The voice map representation clearly visualized and quantified phonatory changes post-treatment for UVFP, with potential applications in clinical monitoring and outcome evaluation.

Ort, förlag, år, upplaga, sidor
Firenze Univ Press, 2025
Serie
Proceedings E Report, ISSN 2704-601X
Nyckelord
vocal fold paralysis, injection laryngoplasty, voice mapping, electroglottography, vocal fold contact
Nationell ämneskategori
Oto-rino-laryngologi
Identifikatorer
urn:nbn:se:kth:diva-379481 (URN)001686443700001 ()
Konferens
14th International Workshop on MODELS AND ANALYSIS OF VOCAL EMISSIONS FOR BIOMEDICAL APPLICATIONS-MAVEBA, DEC 16-17, 2025, Firenze, ITALY
Anmärkning

Part of ISBN 979-12-215-0820-8; 979-12-215-0821-5

QC 20260416

Tillgänglig från: 2026-04-16 Skapad: 2026-04-16 Senast uppdaterad: 2026-04-16Bibliografiskt granskad
Ternström, S. & Pabon, P. (2025). "Voice Range Profile" or "Voice Map"?: On terms, rationales and techniques. In: L. Frassineti, A. Lanatà, C. Manfredi (Ed.), Models and Analysis of Vocal Emissions for Biomedical Applications: 14th International Workshop. Paper presented at 14th MAVEBA Workshop, 16-17 Dec, Florence, Italy (pp. 135-138). Firenze, Italy: Firenze University Press (FUP)
Öppna denna publikation i ny flik eller fönster >>"Voice Range Profile" or "Voice Map"?: On terms, rationales and techniques
2025 (Engelska)Ingår i: Models and Analysis of Vocal Emissions for Biomedical Applications: 14th International Workshop / [ed] L. Frassineti, A. Lanatà, C. Manfredi, Firenze, Italy: Firenze University Press (FUP), 2025, s. 135-138Konferensbidrag, Publicerat paper (Refereegranskat)
Abstract [en]

Let “voice range profile” (a.k.a, “phonetogram”) be the term for a graph of the maximum phonatory range of a voice on the fo×SPL plane, i.e., a closed contour. Let “voice map” be the term for a map of a scalar metric over some relevant range,not necessarily to the extremes, on that same plane, i.e., a 2D scalar field. For imaging several metrics, one voice map can have several “layers”, all derived from the same recording. This paradigm for collection and collation of voice data is useful, because it accounts for how the chosen metrics vary systematically with fo and SPL. Both fo and SPL are influential and typically nonlinear covariates of other voice metrics. Not accounting for them can obscure the effects of an intervention. Here we summarize some central concepts, rationales and techniques related to voice mapping.

Ort, förlag, år, upplaga, sidor
Firenze, Italy: Firenze University Press (FUP), 2025
Serie
Models and Analysis of Vocal Emissions for Biomedical Applications, ISSN 2704-601X, E-ISSN ISSN 2704-5846 ; 139
Nyckelord
voice analysis, voice map, voice range profile, electroglottography
Nationell ämneskategori
Signalbehandling
Forskningsämne
Tal- och musikkommunikation
Identifikatorer
urn:nbn:se:kth:diva-374335 (URN)
Konferens
14th MAVEBA Workshop, 16-17 Dec, Florence, Italy
Projekt
Språkbanken Tal
Forskningsfinansiär
Vetenskapsrådet
Anmärkning

Part of ISBN 9791221508208, 9791221508215

QC 20251218

Tillgänglig från: 2025-12-17 Skapad: 2025-12-17 Senast uppdaterad: 2026-02-25Bibliografiskt granskad
Ternström, S., Bernardoni, N. H., Birkholz, P., Guasch, O. & Gully, A. (Eds.). (2024). Computational Analysis and Simulation of the Human Voice (Dagstuhl Seminar 24242). Paper presented at Dagstuhl Seminar 24242. Schloss Dagstuhl – Leibniz-Zentrum für Informatik, 14(6)
Öppna denna publikation i ny flik eller fönster >>Computational Analysis and Simulation of the Human Voice (Dagstuhl Seminar 24242)
Visa övriga...
2024 (Engelska)Proceedings (redaktörskap) (Övrigt vetenskapligt)
Abstract [en]

This report documents the program and the outcomes of Dagstuhl Seminar 24242 "Computational Analysis and Simulation of the Human Voice", which was held from the 9th to the 14th of June, 2024. The seminar addressed key issues for a better understanding of the human voice by focusing on four main areas: voice analysis, visualisation techniques, simulation methods, and data analysis with machine learning. There has been enormous progress in recent years in all these fields. The seminar brought together a number of experts from fields as diverse as computer science, logopedics and phoniatrics, clinicians, acoustics and audio engineering, electronics, musicology, speech and hearing sciences, physics and mathematics. The schedule was quite flexible, including inspirational talks in the main areas, interactive working groups, sharing of conclusions and discussions, presentation of successes and failures to learn from, and a large number of free talks that emerged throughout the days. The variety of topics and participants created a highly enriching environment from which novel proposals for future research and collaboration emerged, as well as the collective writing of a paper on the state of the art and future perspectives in human voice research.

Ort, förlag, år, upplaga, sidor
Schloss Dagstuhl – Leibniz-Zentrum für Informatik, 2024. s. 24
Serie
Dagstuhl Reports, ISSN 2192-5283 ; 14
Nyckelord
voice analysis, voice simulation, voice visualization
Nationell ämneskategori
Bioinformatik (beräkningsbiologi) Annan data- och informationsvetenskap
Forskningsämne
Tal- och musikkommunikation
Identifikatorer
urn:nbn:se:kth:diva-357969 (URN)10.4230/DagRep.14.6.84 (DOI)
Konferens
Dagstuhl Seminar 24242
Anmärkning

QC 20250113

Tillgänglig från: 2024-12-21 Skapad: 2024-12-21 Senast uppdaterad: 2025-01-13Bibliografiskt granskad
Iob, N. A., He, L., Ternström, S., Cai, H. & Brockmann-Bauser, M. (2024). Effects of Speech Characteristics on Electroglottographic and Instrumental Acoustic Voice Analysis Metrics in Women With Structural Dysphonia Before and After Treatment. Journal of Speech, Language and Hearing Research, 67(6), 1660-1681
Öppna denna publikation i ny flik eller fönster >>Effects of Speech Characteristics on Electroglottographic and Instrumental Acoustic Voice Analysis Metrics in Women With Structural Dysphonia Before and After Treatment
Visa övriga...
2024 (Engelska)Ingår i: Journal of Speech, Language and Hearing Research, ISSN 1092-4388, E-ISSN 1558-9102, Vol. 67, nr 6, s. 1660-1681Artikel i tidskrift (Refereegranskat) Published
Abstract [en]

Purpose: Literature suggests a dependency of the acoustic metrics, smoothed cepstral peak prominence (CPPS) and harmonics-to-noise ratio (HNR), on human voice loudness and fundamental frequency (fo). Even though this has been explained with different oscillatory patterns of the vocal folds, so far, it has not been specifically investigated. In the present work, the influence of three elicitation levels, calibrated sound pressure level (SPL), fo and vowel on the electroglottographic (EGG) and time-differentiated EGG (dEGG) metrics hybrid open quotient (OQ), dEGG OQ and peak dEGG, as well as on the acous-tic metrics CPPS and HNR, was examined, and their suitability for voice assess-ment was evaluated. Method: In a retrospective study, 29 women with a mean age of 25 years (± 8.9, range: 18–53) diagnosed with structural vocal fold pathologies were examined before and after voice therapy or phonosurgery. Both acoustic and EGG signals were recorded simultaneously during the phonation of the sustained vowels /ɑ/, /i/, and /u/ at three elicited levels of loudness (soft/comfortable/loud) and unconstrained fo conditions. Results: A linear mixed-model analysis showed a significant effect of elicitation effort levels on peak dEGG, HNR, and CPPS (all p < .01). Calibrated SPL significantly influenced HNR and CPPS (both p < .01). Furthermore, F0had asignificant effect on peak dEGG and CPPS (p < .0001). All metrics showed significant changes with regard to vowel (all p < .05). However, the treatment had no effect on the examined metrics, regardless of the treatment type (surgery vs. voice therapy). Conclusions: The value of the investigated metrics for voice assessment purposes when sampled without sufficient control of SPL and fo is limited, in that they are significantly influenced by the phonatory context, be it speech or elicited sustained vowels. Future studies should explore the diagnostic value of new data collation approaches such as voice mapping, which take SPL and fo effects into account.

Ort, förlag, år, upplaga, sidor
American Speech Language Hearing Association, 2024
Nationell ämneskategori
Oto-rino-laryngologi
Forskningsämne
Tal- och musikkommunikation
Identifikatorer
urn:nbn:se:kth:diva-346605 (URN)10.1044/2024_JSLHR-23-00253 (DOI)001245110000002 ()38758676 (PubMedID)2-s2.0-85192238446 (Scopus ID)
Anmärkning

QC 20240703

Tillgänglig från: 2024-05-20 Skapad: 2024-05-20 Senast uppdaterad: 2025-02-21Bibliografiskt granskad
Organisationer
Identifikatorer
ORCID-id: ORCID iD iconorcid.org/0000-0002-3362-7518

Sök vidare i DiVA

Visa alla publikationer