kth.sePublications KTH
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
The problems with using STNs to align CNN feature maps
KTH, School of Electrical Engineering and Computer Science (EECS), Computer Science, Computational Science and Technology (CST). (Computational Brain Science)ORCID iD: 0000-0001-8548-5788
KTH, School of Electrical Engineering and Computer Science (EECS), Computer Science, Computational Science and Technology (CST). (Computational Brain Science)ORCID iD: 0000-0003-0011-6444
KTH, School of Electrical Engineering and Computer Science (EECS), Computer Science, Computational Science and Technology (CST). (Computational Brain Science Lab)ORCID iD: 0000-0002-9081-2170
2020 (English)Report (Other academic)
Abstract [en]

Spatial transformer networks (STNs) were designed to enable CNNs to learn invariance to image transformations. STNs were originally proposed to transform CNN feature maps as well as input images. This enables the use of more complex features when predicting transformation parameters. However, since STNs perform a purely spatial transformation, they do not, in the general case, have the ability to align the feature maps of a transformed image and its original. We present a theoretical argument for this and investigate the practical implications, showing that this inability is coupled with decreased classification accuracy. We advocate taking advantage of more complex features in deeper layers by instead sharing parameters between the classification and the localisation network.

Place, publisher, year, edition, pages
2020. , p. 2
National Category
Computer graphics and computer vision
Research subject
Computer Science
Identifiers
URN: urn:nbn:se:kth:diva-363943DOI: 10.48550/arXiv.2001.05858OAI: oai:DiVA.org:kth-363943DiVA, id: diva2:1962018
Funder
Swedish Research Council, 2018-03586
Note

QC 20250814

Available from: 2025-05-28 Created: 2025-05-28 Last updated: 2025-08-14

Open Access in DiVA

fulltext(384 kB)66 downloads
File information
File name FULLTEXT01.pdfFile size 384 kBChecksum SHA-512
c4c5ae5e90b197ddc5c6d74239236a757b3b2b15782e41ca33b998440a737bae7deb8be99253a5e88786e17782f0be4f8d810bbe33f480a834e3831c073519b6
Type fulltextMimetype application/pdf

Other links

Publisher's full text

Authority records

Finnveden, LukasJansson, Ylva

Search in DiVA

By author/editor
Finnveden, LukasJansson, YlvaLindeberg, Tony
By organisation
Computational Science and Technology (CST)
Computer graphics and computer vision

Search outside of DiVA

GoogleGoogle Scholar
Total: 66 downloads
The number of downloads is the sum of all downloads of full texts. It may include eg previous versions that are now no longer available

doi
urn-nbn

Altmetric score

doi
urn-nbn
Total: 332 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf