kth.sePublications KTH
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Inability of spatial transformations of CNN feature maps to support invariant recognition
KTH, School of Electrical Engineering and Computer Science (EECS), Computer Science, Computational Science and Technology (CST). (Computational Brain Science Lab)ORCID iD: 0000-0003-0011-6444
KTH, School of Electrical Engineering and Computer Science (EECS), Computer Science, Computational Science and Technology (CST). (Computational Brain Science Lab)ORCID iD: 0000-0001-8548-5788
KTH, School of Electrical Engineering and Computer Science (EECS), Computer Science, Computational Science and Technology (CST). (Computational Brain Science Lab)ORCID iD: 0000-0002-9081-2170
2020 (English)Report (Other academic)
Abstract [en]

A large number of deep learning architectures use spatial transformations of CNN feature maps or filters to better deal with variability in object appearance caused by natural image transformations. In this paper, we prove that spatial transformations of CNN feature maps cannot align the feature maps of a transformed image to match those of its original, for general affine transformations, unless the extracted features are themselves invariant. Our proof is based on elementary analysis for both the single- and multi-layer network case. The results imply that methods based on spatial transformations of CNN feature maps or filters cannot replace image alignment of the input and cannot enable invariant recognition for general affine transformations, specifically not for scaling transformations or shear transformations. For rotations and reflections, spatially transforming feature maps or filters can enable invariance but only for networks with learnt or hardcoded rotation- or reflection-invariant features

Place, publisher, year, edition, pages
2020. , p. 22
Keywords [en]
deep learning, convolutional neural networks, invariant neural networks, spatial transformer networks
National Category
Computer graphics and computer vision
Identifiers
URN: urn:nbn:se:kth:diva-272970OAI: oai:DiVA.org:kth-272970DiVA, id: diva2:1428472
Funder
Swedish Research Council, 2018-03586
Note

QC 20200507

Available from: 2020-05-05 Created: 2020-05-05 Last updated: 2025-02-07Bibliographically approved

Open Access in DiVA

fulltext(1271 kB)209 downloads
File information
File name FULLTEXT01.pdfFile size 1271 kBChecksum SHA-512
0a8b6fb3c41cdb8efd0bdc45c235c07d45a4eec5a385762dfba146b1d73bdb32b73126d63f62fb7eadf0d32a97758e8820616d000af64a2a76df993970a1b21d
Type fulltextMimetype application/pdf

Other links

arXiv:2004.14716

Authority records

Jansson, YlvaFinnveden, LukasLindeberg, Tony

Search in DiVA

By author/editor
Jansson, YlvaMaydanskiy, MaksimFinnveden, LukasLindeberg, Tony
By organisation
Computational Science and Technology (CST)
Computer graphics and computer vision

Search outside of DiVA

GoogleGoogle Scholar
Total: 209 downloads
The number of downloads is the sum of all downloads of full texts. It may include eg previous versions that are now no longer available

urn-nbn

Altmetric score

urn-nbn
Total: 593 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf