kth.sePublications KTH
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Scale generalisation properties of extended scale-covariant and scale-invariant Gaussian derivative networks on image datasets with spatial scaling variations
KTH, School of Electrical Engineering and Computer Science (EECS), Computer Science, Computational Science and Technology (CST).ORCID iD: 0009-0004-7143-5447
KTH, School of Electrical Engineering and Computer Science (EECS), Computer Science, Computational Science and Technology (CST).ORCID iD: 0000-0002-9081-2170
2025 (English)In: Journal of Mathematical Imaging and Vision, ISSN 0924-9907, E-ISSN 1573-7683, Vol. 67, no 3, p. 29:1-29:39, article id 29Article in journal (Refereed) Published
Abstract [en]

Due to the variabilities in image structures caused by perspective scaling transformations, it is essential for deep networks to have an ability to generalise to scales not seen during training. This paper presents an in-depth analysis of the scale generalisation properties of the scale-covariant and scale-invariant Gaussian derivative networks, complemented with both conceptual and algorithmic extensions. For this purpose, Gaussian derivative networks (GaussDerNets) are evaluated on new rescaled versions of the Fashion-MNIST and the CIFAR-10 datasets, with spatial scaling variations over a factor of 4 in the testing data, that are not present in the training data. Additionally, evaluations on the previously existing STIR datasets show that the GaussDerNets achieve better scale generalisation than previously reported for these datasets for other types of deep networks. We first experimentally demonstrate that the GaussDerNets have quite good scale generalisation properties on the new datasets and that average pooling of feature responses over scales may sometimes also lead to better results than the previously used approach of max pooling over scales. Then, we demonstrate that using a spatial max pooling mechanism after the final layer enables localisation of non-centred objects in the image domain, with maintained scale generalisation properties. We also show that regularisation during training, by applying dropout across the scale channels, referred to as scale-channel dropout, improves both the performance and the scale generalisation. In additional ablation studies, we show that, for the rescaled CIFAR-10 dataset, basing the layers in the GaussDerNets on derivatives up to order three leads to better performance and scale generalisation for coarser scales, whereas networks based on derivatives up to order two achieve better scale generalisation for finer scales. Moreover, we demonstrate that discretisations of GaussDerNets based on the discrete analogue of the Gaussian kernel in combination with central difference operators perform best or among the best, compared to a set of other discrete approximations of the Gaussian derivative kernels. Furthermore, we show that the improvement in performance obtained by learning the scale values of the Gaussian derivatives, as opposed to using the previously proposed choice of a fixed logarithmic distribution of the scale levels, is usually only minor, thus supporting the previously postulated choice of using a logarithmic distribution as a very reasonable prior. Finally, by visualising the activation maps and the learned receptive fields, we demonstrate that the GaussDerNets have very good explainability properties.

Place, publisher, year, edition, pages
Springer Nature , 2025. Vol. 67, no 3, p. 29:1-29:39, article id 29
Keywords [en]
Deep learning, Gaussian derivative, Receptive fields, Scale covariance, Scale generalisation, Scale invariance, Scale selection, Scale space
National Category
Computer graphics and computer vision
Research subject
Computer Science
Identifiers
URN: urn:nbn:se:kth:diva-363792DOI: 10.1007/s10851-025-01245-xISI: 001488261900001Scopus ID: 2-s2.0-105004740771OAI: oai:DiVA.org:kth-363792DiVA, id: diva2:1959888
Projects
Covariant and invariant deep networks
Funder
Swedish Research Council, 2018-03586, 2022-02969
Note

QC 20250523

Ej dubblett

Available from: 2025-05-21 Created: 2025-05-21 Last updated: 2025-10-27Bibliographically approved

Open Access in DiVA

fulltext(3801 kB)77 downloads
File information
File name FULLTEXT01.pdfFile size 3801 kBChecksum SHA-512
ba13e8c577c93a6292e52dd627a41be05a4eee152ad43e4f74cb3fbf0ae3e403cde84969b797cbca9043dee0b6253a998d8910bf32a09a890f1debe962aa7906
Type fulltextMimetype application/pdf
Supplement(482 kB)39 downloads
File information
File name FULLTEXT02.pdfFile size 482 kBChecksum SHA-512
b715331168528a21f84850ac4056610d07e0553188f70ba9780f59e6a6e0efac8c98e2879554329d464630cf71259a3d94bc9847e29421745ebee08ef8722c95
Type fulltextMimetype application/pdf

Other links

Publisher's full textScopus

Authority records

Perzanowski, AndrzejLindeberg, Tony

Search in DiVA

By author/editor
Perzanowski, AndrzejLindeberg, Tony
By organisation
Computational Science and Technology (CST)
In the same journal
Journal of Mathematical Imaging and Vision
Computer graphics and computer vision

Search outside of DiVA

GoogleGoogle Scholar
Total: 116 downloads
The number of downloads is the sum of all downloads of full texts. It may include eg previous versions that are now no longer available

doi
urn-nbn

Altmetric score

doi
urn-nbn
Total: 1142 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf