kth.sePublications KTH
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Understanding when spatial transformer networks do not support invariance, and what to do about it
KTH, School of Electrical Engineering and Computer Science (EECS), Computer Science, Computational Science and Technology (CST). (Computational Brain Science Lab)ORCID iD: 0000-0001-8548-5788
KTH, School of Electrical Engineering and Computer Science (EECS), Computer Science, Computational Science and Technology (CST). (Computational Brain Science Lab)ORCID iD: 0000-0003-0011-6444
KTH, School of Electrical Engineering and Computer Science (EECS), Computer Science, Computational Science and Technology (CST). (Computational Brain Science Lab)ORCID iD: 0000-0002-9081-2170
2020 (English)Report (Other academic)
Abstract [en]

Spatial transformer networks (STNs) were designed to enable convolutional neural networks (CNNs) to learn invariance to image transformations. STNs were originally proposed to transform CNN feature maps as well as input images. This enables the use of more complex features when predicting transformation parameters. However, since STNs perform a purely spatial transformation, they do not, in the general case, have the ability to align the feature maps of a transformed image with those of its original. STNs are therefore unable to support invariance when transforming CNN feature maps. We present a simple proof for this and study the practical implications, showing that this inability is coupled with decreased classification accuracy. We therefore investigate alternative STN architectures that make use of complex features. We find that while deeper localization networks are difficult to train, localization networks that share parameters with the classification network remain stable as they grow deeper, which allows for higher classification accuracy on difficult datasets. Finally, we explore the interaction between localization network complexity and iterative image alignment.

Place, publisher, year, edition, pages
2020. , p. 12
Keywords [en]
deep learning, convolutional neural networks, invariant neural networks, spatial transformer networks
National Category
Computer graphics and computer vision
Research subject
Computer Science
Identifiers
URN: urn:nbn:se:kth:diva-272971OAI: oai:DiVA.org:kth-272971DiVA, id: diva2:1428271
Funder
Swedish Research Council, 2018-03586
Note

Not duplicate with DiVA 1516191

QC 20200511

Available from: 2020-05-05 Created: 2020-05-05 Last updated: 2025-02-07Bibliographically approved

Open Access in DiVA

fulltext(1275 kB)2497 downloads
File information
File name FULLTEXT06.pdfFile size 1275 kBChecksum SHA-512
54b8ba8e2d87a7aa5eec2d1152680617ec92f1147f20110ee1eda66c7b152feb0aaffd4b318f7b6e2bca602b52122338ebb6daac2cb1e19abe34d15f133b53e9
Type fulltextMimetype application/pdf

Other links

arXiv:2004.11678

Authority records

Finnveden, LukasJansson, YlvaLindeberg, Tony

Search in DiVA

By author/editor
Finnveden, LukasJansson, YlvaLindeberg, Tony
By organisation
Computational Science and Technology (CST)
Computer graphics and computer vision

Search outside of DiVA

GoogleGoogle Scholar
Total: 3095 downloads
The number of downloads is the sum of all downloads of full texts. It may include eg previous versions that are now no longer available

urn-nbn

Altmetric score

urn-nbn
Total: 1849 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf