kth.sePublications KTH
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
SpaNN: Detecting Multiple Adversarial Patches on CNNs by Spanning Saliency Thresholds
KTH, School of Electrical Engineering and Computer Science (EECS), Computer Science, Network and Systems Engineering.ORCID iD: 0009-0009-8875-8329
KTH, School of Electrical Engineering and Computer Science (EECS), Computer Science, Network and Systems Engineering.ORCID iD: 0000-0002-4876-0223
KTH, School of Electrical Engineering and Computer Science (EECS), Intelligent systems, Decision and Control Systems (Automatic Control).ORCID iD: 0000-0003-1835-2963
2025 (English)In: Proceedings - 2025 IEEE Conference on Secure and Trustworthy Machine Learning, SaTML 2025, Institute of Electrical and Electronics Engineers (IEEE) , 2025, p. 459-478Conference paper, Published paper (Refereed)
Abstract [en]

State-of-the-art convolutional neural network models for object detection and image classification are vulnerable to physically realizable adversarial perturbations, such as patch attacks. Existing defenses have focused, implicitly or explicitly, on single-patch attacks, leaving their sensitivity to the number of patches as an open question or rendering them computationally infeasible or inefficient against attacks consisting of multiple patches in the worst cases. In this work, we propose SpaNN, an attack detector whose computational complexity is independent of the expected number of adversarial patches. The key novelty of the proposed detector is that it builds an ensemble of binarized feature maps by applying a set of saliency thresholds to the neural activations of the first convolutional layer of the victim model. It then performs clustering on the ensemble and uses the cluster features as the input to a classifier for attack detection. Contrary to existing detectors, SpaNN does not rely on a fixed saliency threshold for identifying adversarial regions, which makes it robust against white box adversarial attacks. We evaluate SpaNN on four widely used data sets for object detection and classification, and our results show that SpaNN outperforms state-of-the-art defenses by up to 11 and 27 percentage points in the case of object detection and the case of image classification, respectively. Our code is available at https://github.com/gerkbyrd/SpaNN.

Place, publisher, year, edition, pages
Institute of Electrical and Electronics Engineers (IEEE) , 2025. p. 459-478
Keywords [en]
adversarial machine learning, adversarial patch attacks, Convolutional neural networks
National Category
Computer graphics and computer vision Signal Processing
Identifiers
URN: urn:nbn:se:kth:diva-364401DOI: 10.1109/SaTML64287.2025.00032ISI: 001511726400024Scopus ID: 2-s2.0-105007307138OAI: oai:DiVA.org:kth-364401DiVA, id: diva2:1968215
Conference
2025 IEEE Conference on Secure and Trustworthy Machine Learning, SaTML 2025, Copenhagen, Denmark, Apr 9 2025 - Apr 11 2025
Note

Part of ISBN 9798331517113

QC 20250613

Available from: 2025-06-12 Created: 2025-06-12 Last updated: 2025-12-08Bibliographically approved

Open Access in DiVA

No full text in DiVA

Other links

Publisher's full textScopus

Authority records

Byrd Victorica, MauricioDán, GyörgySandberg, Henrik

Search in DiVA

By author/editor
Byrd Victorica, MauricioDán, GyörgySandberg, Henrik
By organisation
Network and Systems EngineeringDecision and Control Systems (Automatic Control)
Computer graphics and computer visionSignal Processing

Search outside of DiVA

GoogleGoogle Scholar

doi
urn-nbn

Altmetric score

doi
urn-nbn
Total: 158 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf