kth.sePublications KTH
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Exploring the ability of CNNs to generalise to previously unseen scales over wide scale ranges
KTH, School of Electrical Engineering and Computer Science (EECS), Computer Science, Computational Science and Technology (CST). (Computational Brain Science Lab)ORCID iD: 0000-0003-0011-6444
KTH, School of Electrical Engineering and Computer Science (EECS), Computer Science, Computational Science and Technology (CST). (Computational Brain Science Lab)ORCID iD: 0000-0002-9081-2170
2020 (English)Report (Other academic)
Abstract [en]

The ability to handle large scale variations is crucial for many real world visual tasks. A straightforward approach for handling scale in a deep network is to process an image at several scales simultaneously in a set of scale channels. Scale invariance can then, in principle, be achieved by using weight sharing between the scale channels together with max or average pooling over the outputs from the scale channels. The ability of such scale channel networks to generalise to scales not present in the training set over significant scale ranges has, however, not previously been explored. We, therefore, present a theoretical analysis of invariance and covariance properties of scale channel networks and perform an experimental evaluation of the ability of different types of scale channel networks to generalise to previously unseen scales. We identify limitations of previous approaches and propose a new type of foveated scale channel architecture, where the scale channels process increasingly larger parts of the image with decreasing resolution. Our proposed FovMax and FovAvg networks perform almost identically over a scale range of 8 also when training on single scale training data and give improvements in the small sample regime.

Place, publisher, year, edition, pages
2020.
Keywords [en]
deep learning, convolutional neural networks, invariant neural networks, scale invariance
National Category
Computer graphics and computer vision
Identifiers
URN: urn:nbn:se:kth:diva-272013OAI: oai:DiVA.org:kth-272013DiVA, id: diva2:1423788
Funder
Swedish Research Council, 2018-03586
Note

Not duplicate with 1515273

QC 20220517

Available from: 2020-04-15 Created: 2020-04-15 Last updated: 2025-02-07Bibliographically approved

Open Access in DiVA

fulltext(4212 kB)266 downloads
File information
File name FULLTEXT06.pdfFile size 4212 kBChecksum SHA-512
9c88027d0b9e0887c3d8cfe9115fe708ff72ab8bc7d99a82f553e491317f0df3c8c61d7eaba045f45956bf6e7d916f3b7a2e12b16b1439d9ee66dbbb7dc38bde
Type fulltextMimetype application/pdf

Other links

arXiv preprint arXiv:2004.01536https://arxiv.org/abs/2004.01536

Authority records

Jansson, YlvaLindeberg, Tony

Search in DiVA

By author/editor
Jansson, YlvaLindeberg, Tony
By organisation
Computational Science and Technology (CST)
Computer graphics and computer vision

Search outside of DiVA

GoogleGoogle Scholar
Total: 310 downloads
The number of downloads is the sum of all downloads of full texts. It may include eg previous versions that are now no longer available

urn-nbn

Altmetric score

urn-nbn
Total: 1179 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf