kth.sePublications KTH
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Tuning frequency bias of state space models
Center for Applied Mathematics, Cornell University, Ithaca, NY 14853, USA, United States.
Data Science Institute, University of Chicago, Chicago, IL 60637, USA, United States.
KTH, School of Engineering Sciences (SCI), Mathematics (Dept.), Probability, Mathematical Physics and Statistics. KTH, Centres, Nordic Institute for Theoretical Physics NORDITA.ORCID iD: 0000-0002-4649-673X
Lawrence Berkeley National Laboratory, Berkeley, CA 94720, USA, United States; Department of Statistics, University of California at Berkeley, Berkeley, CA 94720, USA, United States.
Show others and affiliations
2025 (English)In: 13th International Conference on Learning Representations, ICLR 2025, International Conference on Learning Representations, ICLR , 2025, p. 92498-92521Conference paper, Published paper (Refereed)
Abstract [en]

State space models (SSMs) leverage linear, time-invariant (LTI) systems to effectively learn sequences with long-range dependencies. By analyzing the transfer functions of LTI systems, we find that SSMs exhibit an implicit bias toward capturing low-frequency components more effectively than high-frequency ones. This behavior aligns with the broader notion of frequency bias in deep learning model training. We show that the initialization of an SSM assigns it an innate frequency bias and that training the model in a conventional way does not alter this bias. Based on our theory, we propose two mechanisms to tune frequency bias: either by scaling the initialization to tune the inborn frequency bias; or by applying a Sobolev-norm-based filter to adjust the sensitivity of the gradients to high-frequency inputs, which allows us to change the frequency bias via training. Using an image-denoising task, we empirically show that we can strengthen, weaken, or even reverse the frequency bias using both mechanisms. By tuning the frequency bias, we can also improve SSMs' performance on learning long-range sequences, averaging an 88.26% accuracy on the Long-Range Arena (LRA) benchmark tasks.

Place, publisher, year, edition, pages
International Conference on Learning Representations, ICLR , 2025. p. 92498-92521
National Category
Control Engineering
Identifiers
URN: urn:nbn:se:kth:diva-385674Scopus ID: 2-s2.0-105010273425OAI: oai:DiVA.org:kth-385674DiVA, id: diva2:2087564
Conference
13th International Conference on Learning Representations, ICLR 2025, Singapore, Singapore, Apr 24 2025 - Apr 28 2025
Note

Part of ISBN 9798331320850

QC 20260721

Available from: 2026-07-21 Created: 2026-07-21 Last updated: 2026-07-21Bibliographically approved

Open Access in DiVA

No full text in DiVA

Scopus

Authority records

Lim, Soon Hoe

Search in DiVA

By author/editor
Lim, Soon Hoe
By organisation
Probability, Mathematical Physics and StatisticsNordic Institute for Theoretical Physics NORDITA
Control Engineering

Search outside of DiVA

GoogleGoogle Scholar

urn-nbn

Altmetric score

urn-nbn
Total: 1 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf