kth.sePublications KTH
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Neural synthesis of sound effects using flow-based deep generative models
KTH, School of Engineering Sciences (SCI), Mathematics (Dept.).
2022 (English)Independent thesis Advanced level (degree of Master (Two Years)), 20 credits / 30 HE creditsStudent thesisAlternative title
Neural syntes av ljudeffekter med hjälp av flödesbaserade djupgenerativa modeller (Swedish)
Abstract [en]

Generating diverse sound effects for video games is a consuming task that grows with the size and complexity of the games. We adopt WaveFlow, a flow-based deep generative model intended for speech synthesis, to generate variations of explosion sounds, given lower-dimensional conditioners in the form of mel spectrograms.This work suggests that it is possible to use flow-based models to generate high-quality raw audio waveforms of sound effects, learning from small datasets, and it proposes metrics for evaluating the quality of the generated audio. It is, to the best of our knowledge, the first adaptation of these techniques to sound effects generation.

Abstract [sv]

Att skapa olika ljudeffekter för videospel är en krävande uppgift som växer med spelens storlek och komplexitet. Vi använder WaveFlow, en flödesbaserad djup generativ modell som är avsedd för talsyntes, för att generera variationer av explosionsljud med hjälp av lägre dimensionella villkor i form av mel-spektrogram.Det här arbetet visar att det är möjligt att använda flödesbaserade modeller för att generera råa ljudvågformer av hög kvalitet för ljudeffekter, inlärning från små datamängder, och föreslår mätvärden för att utvärdera kvaliteten på det genererade ljudet. Det är, såvitt vi vet, den första anpassningen av dessa tekniker till generering av ljudeffekter.

Place, publisher, year, edition, pages
2022.
Series
TRITA-SCI-GRU ; 2022:377
Keywords [en]
Machine Learning, Applied Mathematics, Deep Generative Models, flow models, Signal Processing
Keywords [sv]
Statistik, tillämpad matematik, Machine Learning
National Category
Mathematical sciences
Identifiers
URN: urn:nbn:se:kth:diva-323076OAI: oai:DiVA.org:kth-323076DiVA, id: diva2:1727122
External cooperation
SEED (Electronic Arts)
Subject / course
Mathematics
Educational program
Master of Science - Computer Simulation for Science and Engineering
Supervisors
Examiners
Available from: 2023-02-02 Created: 2023-01-15 Last updated: 2025-10-22Bibliographically approved

Open Access in DiVA

fulltext(10862 kB)166 downloads
File information
File name FULLTEXT02.pdfFile size 10862 kBChecksum SHA-512
a20a26ba24714485113c002168f2b3530176ae2fd2ea38f2294a5f8119d206331fe124a88321f67a4503ea6df00ee03fb7ea0a3c36b3a91aa76593b4f3eab810
Type fulltextMimetype application/pdf

By organisation
Mathematics (Dept.)
Mathematical sciences

Search outside of DiVA

GoogleGoogle Scholar
Total: 1003 downloads
The number of downloads is the sum of all downloads of full texts. It may include eg previous versions that are now no longer available

urn-nbn

Altmetric score

urn-nbn
Total: 442 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf