kth.sePublications KTH
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Vocal-tract length estimation from vowel formants benchmarked against acoustic pharyngometry
Department of Computational Linguistics, University of Zurich, 8050 Zurich, Switzerland; Zurich Forensic Science Institute, 8004 Zurich, Switzerland.ORCID iD: 0000-0003-1102-4916
Department of Computational Linguistics, University of Zurich, 8050 Zurich, Switzerland.
KTH, School of Electrical Engineering and Computer Science (EECS), Speech, Music and Hearing. Centre for Cultural Evolution, Department of Psychology, Stockholm University, 106 91 Stockholm, Sweden.ORCID iD: 0000-0002-6739-0838
Department of Computational Linguistics, University of Zurich, 8050 Zurich, Switzerland.ORCID iD: 0000-0002-8494-6025
Show others and affiliations
2026 (English)In: Journal of the Acoustical Society of America, ISSN 0001-4966, E-ISSN 1520-8524, Vol. 159, no 6, p. 5650-5666Article in journal (Refereed) Published
Abstract [en]

Estimating vocal-tract length (VTL) from vowel formants can aid speaker normalization, but few methods have been benchmarked against an anatomical reference in the same speakers. We combined acoustic pharyngometry (APh) and speech data from 42 adults to benchmark eight widely used formant-based VTL estimators against incisors-to-glottis length and to test an interpretable two-stage bias-corrected linear estimator. Across more than 400 000 central frames with valid F 1–F 4, traditional quarter-wave, odd-harmonic, and dispersion-type estimators correlated with VTL APh but showed poor out-of-sample anatomical recovery and strong calibration compression. Re-estimated one-stage linear models reduced mean absolute error (MAE; median ≈1.0 cm) but still overestimated shorter tracts and underestimated longer tracts. A two-stage model markedly improved calibration and agreement, outperforming one-stage linear and nonlinear alternatives (median per-vowel MAE 0.39 cm, median out-of-sample R 2 = 0.83). Front and front-rounded vowels were the most informative. Speaker-level 95% limits of agreement were about ±0.9 cm, indicating that the method is better suited to aggregated tract-scale estimation than to direct anatomical measurement. These results identify calibration bias as a central limitation of standard formant-based VTL estimators and provide a practical, interpretable route to tract-scale estimation from similarly processed labeled-vowel data under matched conditions.

Place, publisher, year, edition, pages
Acoustical Society of America (ASA) , 2026. Vol. 159, no 6, p. 5650-5666
National Category
Probability Theory and Statistics
Identifiers
URN: urn:nbn:se:kth:diva-385339DOI: 10.1121/10.0044192ISI: 001798874500001PubMedID: 42329038Scopus ID: 2-s2.0-105042720490OAI: oai:DiVA.org:kth-385339DiVA, id: diva2:2086251
Note

QC 20260713

Available from: 2026-07-13 Created: 2026-07-13 Last updated: 2026-07-13Bibliographically approved

Open Access in DiVA

No full text in DiVA

Other links

Publisher's full textPubMedScopus

Authority records

Ekström, Axel G.

Search in DiVA

By author/editor
Friedrichs, DanielEkström, Axel G.Dellwo, VolkerMoran, Steven
By organisation
Speech, Music and Hearing
In the same journal
Journal of the Acoustical Society of America
Probability Theory and Statistics

Search outside of DiVA

GoogleGoogle Scholar

doi
pubmed
urn-nbn

Altmetric score

doi
pubmed
urn-nbn
Total: 5 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf