kth.sePublications KTH
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Scale-invariant scale-channel networks: Deep networks that generalise to previously unseen scales
KTH, School of Electrical Engineering and Computer Science (EECS), Computer Science, Computational Science and Technology (CST). (Computational Brain Science Lab)ORCID iD: 0000-0003-0011-6444
KTH, School of Electrical Engineering and Computer Science (EECS), Computer Science, Computational Science and Technology (CST). (Computational Brain Science Lab)ORCID iD: 0000-0002-9081-2170
2022 (English)In: Journal of Mathematical Imaging and Vision, ISSN 0924-9907, E-ISSN 1573-7683, Vol. 64, no 5, p. 506-536Article in journal (Refereed) Published
Abstract [en]

The ability to handle large scale variations is crucial for many real world visual tasks. A straightforward approach for handling scale in a deep network is toprocess an image at several scales simultaneously in a set of scale channels. Scale invariance can then, in principle, be achieved by using weight sharing between the scale channels together with max or average pooling over the outputs from the scale channels. The ability of such scale-channel networks to generalise to scales not present in the training set over significant scale ranges has, however, not previously been explored. 

In this paper, we present a systematic study of this methodology by implementing different types of scale-channel networks and evaluating their ability to generalise to previously unseen scales. We develop a formalism for analysing the covariance and invariance properties of scale-channel networks, including exploring their relations to scale-space theory, and exploring how different design choices, unique to scaling transformations, affect the overall performance of scale-channel networks. We first show that two previously proposed scale-channel network designs, in one case, generalise no better than a standard CNN to scales not present in the training set, and in the second case, have limited scale generalisation ability. We explain theoretically and demonstrate experimentally why generalisation fails or is limited in these cases. We then propose a new type of foveated scale-channel architecture, where the scale channels process increasingly larger parts of the image with decreasing resolution. This new type of scale-channel network is shown to generalise extremely well, provided sufficient image resolution and the absence of boundary effects. Our proposed FovMax and FovAvg networks perform almost identically over a scale range of 8, also when training on single-scale training data, and do also give improved performance  when learning from datasets with large scale variations in the small sample regime.

Place, publisher, year, edition, pages
Springer Science+Business Media B.V., 2022. Vol. 64, no 5, p. 506-536
Keywords [en]
Deep learning, Convolutional neural networks, Invariant neural networks, Scale covariance, Scale invariance, Scale generalisation, Scale space
National Category
Computer graphics and computer vision
Research subject
Computer Science
Identifiers
URN: urn:nbn:se:kth:diva-309251DOI: 10.1007/s10851-022-01082-2ISI: 000780653200001Scopus ID: 2-s2.0-85127949060OAI: oai:DiVA.org:kth-309251DiVA, id: diva2:1640423
Projects
Scale-space theory for covariant and invariant visual perception
Funder
Swedish Research Council, 2018-03586
Note

QC 20220530

Available from: 2022-02-24 Created: 2022-02-24 Last updated: 2025-02-07Bibliographically approved

Open Access in DiVA

fulltext(2234 kB)287 downloads
File information
File name FULLTEXT02.pdfFile size 2234 kBChecksum SHA-512
d41d1ac7ed52bb703bef0a2ea483c125d6dcbd59a7a00cb618dad935bc957cd8da1836462a602f356e7b3d4670e6b96eb98476217ac00656e3765f15dd19494a
Type fulltextMimetype application/pdf

Other links

Publisher's full textScopus

Authority records

Jansson, YlvaLindeberg, Tony

Search in DiVA

By author/editor
Jansson, YlvaLindeberg, Tony
By organisation
Computational Science and Technology (CST)
In the same journal
Journal of Mathematical Imaging and Vision
Computer graphics and computer vision

Search outside of DiVA

GoogleGoogle Scholar
Total: 294 downloads
The number of downloads is the sum of all downloads of full texts. It may include eg previous versions that are now no longer available

doi
urn-nbn

Altmetric score

doi
urn-nbn
Total: 2066 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf