kth.sePublications KTH
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
GAINS: Human Gaussian Animation Synthesis from Input Speech
KTH, School of Electrical Engineering and Computer Science (EECS).
2025 (English)Independent thesis Advanced level (degree of Master (Two Years)), 20 credits / 30 HE creditsStudent thesisAlternative title
GAINS : Syntes av mänsklig gaussisk animation från talinmatning (Swedish)
Abstract [en]

Generating talking human models from speech input has a long research history, involving advancements in both motion synthesis and audio-visual alignment. Meanwhile, 3D Gaussian Splatting has recently emerged as a powerful rendering technique for modeling complex scenes with high visual fidelity. With more and more research attention, applying 3D Gaussian Splatting to human modeling has now become increasingly feasible in real-time scenarios. This thesis proposes GAINS : an end-to-end method that directly generates animated 3D human models from natural speech input, using 3D Gaussian Splatting as representation method. Unlike traditional multi-stage pipelines that typically rely on third-party tools for video processing, motion capture, or manual annotation, the proposed approach predicts SMPL-based motion directly from speech and synthesizes the corresponding 3D human motions. This significantly reduces system complexity while maintaining high rendering quality and motion fidelity. Through comprehensive evaluations using PSNR, SSIM, and LPIPS metrics against recent methods that rely on rendering frameworks like Pytorch3D,Pyrender,Blender and Nerf, experiments demonstrate that GAINS produces higher-quality animations with superior visual fidelity and greater perceived realism compared to recent multi-stage pipelines User study evaluations with 73 participants also confirm that GAINS human motions scored consistently higher for perceived realism and visual quality compared to conventional multi-stage generation pipelines. This work is among the first to directly generate speech-driven talking human motion using 3D Gaussian point representations, achieving real-time rendering performance.

Abstract [sv]

Att generera pratande 3D-människomodeller från talinmatning har en lång forskningshistoria, med framsteg inom både rörelsesyntes och ljud-bild-synkronisering. Samtidigt har 3D Gaussian Splatting nyligen framträtt som en kraftfull renderingsteknik för modellering av komplexa scener med hög visuell detaljrikedom. Med allt större forskningsintresse har det blivit alltmer genomförbart att tillämpa 3D Gaussian Splatting på människomodellering i realtidsscenarier. Denna avhandling föreslår GAINS: en end-to-end-metod som direkt genererar animerade 3D-människomodeller från naturlig talinmatning, med 3D Gaussian Splatting som representationsmetod. Till skillnad från traditionella flerstegspipelines som ofta förlitar sig på tredjepartsverktyg för videobearbetning, rörelsefångst eller manuell annotering, förutsäger den föreslagna metoden SMPL-baserad rörelse direkt från tal och syntetiserar motsvarande 3D-rörelser. Detta minskar systemkomplexiteten avsevärt samtidigt som hög renderingskvalitet och rörelsetrogenhet bibehålls. Genom omfattande utvärderingar med PSNR-, SSIM- och LPIPS-mått mot nyligen publicerade metoder som använder renderingramverk som Pytorch3D, Pyrender, Blender och NeRF visar experimenten att GAINS producerar animationer med högre kvalitet, överlägsen visuell detaljrikedom och större upplevd realism än flerstegspipelines. Användarstudie med 73 deltagare bekräftar också att GAINS-rörelser konsekvent fick högre betyg för upplevd realism och visuell kvalitet jämfört med konventionella flerstegsmetoder. Detta arbete är bland de första som direkt genererar taldrivna pratande människorörelser med hjälp av 3D-Gaussiska punktrepresentationer, och uppnår realtidsprestanda vid rendering.

Place, publisher, year, edition, pages
2025. , p. 61
Series
TRITA-EECS-EX ; 2025:479
Keywords [en]
Speech‐driven animation, Neural rendering, Real-time rendering, Human motion synthesis, Perceptual evaluation, Audio-visual synthesis
Keywords [sv]
taldriven animation, neural rendering, realtidsrendering, människors rörelsesyntes, perceptuell utvärdering, ljud-bild-syntes
National Category
Computer and Information Sciences
Identifiers
URN: urn:nbn:se:kth:diva-368115OAI: oai:DiVA.org:kth-368115DiVA, id: diva2:1987334
Supervisors
Examiners
Available from: 2025-08-11 Created: 2025-08-05 Last updated: 2025-08-11Bibliographically approved

Open Access in DiVA

fulltext(14807 kB)530 downloads
File information
File name FULLTEXT01.pdfFile size 14807 kBChecksum SHA-512
f1191906bd384271ff86f072c002314799bc079e060e86f1a6aa5e9338ca342be0ecddf7fa3e27946bd8239b2570629d3d029eff2be94e42d3ef467351eb6a40
Type fulltextMimetype application/pdf

By organisation
School of Electrical Engineering and Computer Science (EECS)
Computer and Information Sciences

Search outside of DiVA

GoogleGoogle Scholar
Total: 531 downloads
The number of downloads is the sum of all downloads of full texts. It may include eg previous versions that are now no longer available

urn-nbn

Altmetric score

urn-nbn
Total: 817 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf