Tuning frequency bias of state space modelsShow others and affiliations
2025 (English)In: 13th International Conference on Learning Representations, ICLR 2025, International Conference on Learning Representations, ICLR , 2025, p. 92498-92521Conference paper, Published paper (Refereed)
Abstract [en]
State space models (SSMs) leverage linear, time-invariant (LTI) systems to effectively learn sequences with long-range dependencies. By analyzing the transfer functions of LTI systems, we find that SSMs exhibit an implicit bias toward capturing low-frequency components more effectively than high-frequency ones. This behavior aligns with the broader notion of frequency bias in deep learning model training. We show that the initialization of an SSM assigns it an innate frequency bias and that training the model in a conventional way does not alter this bias. Based on our theory, we propose two mechanisms to tune frequency bias: either by scaling the initialization to tune the inborn frequency bias; or by applying a Sobolev-norm-based filter to adjust the sensitivity of the gradients to high-frequency inputs, which allows us to change the frequency bias via training. Using an image-denoising task, we empirically show that we can strengthen, weaken, or even reverse the frequency bias using both mechanisms. By tuning the frequency bias, we can also improve SSMs' performance on learning long-range sequences, averaging an 88.26% accuracy on the Long-Range Arena (LRA) benchmark tasks.
Place, publisher, year, edition, pages
International Conference on Learning Representations, ICLR , 2025. p. 92498-92521
National Category
Control Engineering
Identifiers
URN: urn:nbn:se:kth:diva-385674Scopus ID: 2-s2.0-105010273425OAI: oai:DiVA.org:kth-385674DiVA, id: diva2:2087564
Conference
13th International Conference on Learning Representations, ICLR 2025, Singapore, Singapore, Apr 24 2025 - Apr 28 2025
Note
Part of ISBN 9798331320850
QC 20260721
2026-07-212026-07-212026-07-21Bibliographically approved