kth.sePublications KTH
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Hierarchical Residual Learning Based Vector Quantized Variational Autoencorder for Image Reconstruction and Generation
KTH, School of Electrical Engineering and Computer Science (EECS), Intelligent systems, Speech, Music and Hearing, TMH, Speech Communication and Technology. Norwegian University of Science and Technology Trondheim, Norway.ORCID iD: 0000-0002-3323-5311
2022 (English)In: The 33rd British Machine Vision Conference Proceedings, 2022Conference paper, Published paper (Refereed)
Abstract [en]

We propose a multi-layer variational autoencoder method, we call HR-VQVAE, thatlearns hierarchical discrete representations of the data. By utilizing a novel objectivefunction, each layer in HR-VQVAE learns a discrete representation of the residual fromprevious layers through a vector quantized encoder. Furthermore, the representations ateach layer are hierarchically linked to those at previous layers. We evaluate our methodon the tasks of image reconstruction and generation. Experimental results demonstratethat the discrete representations learned by HR-VQVAE enable the decoder to reconstructhigh-quality images with less distortion than the baseline methods, namely VQVAE andVQVAE-2. HR-VQVAE can also generate high-quality and diverse images that outperform state-of-the-art generative models, providing further verification of the efficiency ofthe learned representations. The hierarchical nature of HR-VQVAE i) reduces the decoding search time, making the method particularly suitable for high-load tasks and ii) allowsto increase the codebook size without incurring the codebook collapse problem.

Place, publisher, year, edition, pages
2022.
National Category
Computer Sciences
Research subject
Computer Science
Identifiers
URN: urn:nbn:se:kth:diva-324770OAI: oai:DiVA.org:kth-324770DiVA, id: diva2:1743481
Conference
33rd British Machine Vision Conference
Note

QC 20230320

Available from: 2023-03-15 Created: 2023-03-15 Last updated: 2023-03-20Bibliographically approved

Open Access in DiVA

fulltext(17012 kB)1092 downloads
File information
File name FULLTEXT01.pdfFile size 17012 kBChecksum SHA-512
a6980808fb8b0835c2aedaa0292ba2edb7fd7174397b97bc6ecf0f00fac774118bfe21ee85fe4b314b6202356f24a779d41194d900fc554fcc6d86e594aa877e
Type fulltextMimetype application/pdf

Other links

Conference webpage

Authority records

Salvi, Giampiero

Search in DiVA

By author/editor
Salvi, Giampiero
By organisation
Speech Communication and Technology
Computer Sciences

Search outside of DiVA

GoogleGoogle Scholar
Total: 1094 downloads
The number of downloads is the sum of all downloads of full texts. It may include eg previous versions that are now no longer available

urn-nbn

Altmetric score

urn-nbn
Total: 1026 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf