kth.sePublications KTH
Change search
Link to record
Permanent link

Direct link
Publications (10 of 74) Show all publications
Vinuesa, R., Cinnella, P., Rabault, J., Azizpour, H., Bauer, S., Brunton, B. W., . . . Brunton, S. L. (2026). Decoding complexity through machine learning is redefining scientific discovery. Communications Physics, 9(1), Article ID 168.
Open this publication in new window or tab >>Decoding complexity through machine learning is redefining scientific discovery
Show others...
2026 (English)In: Communications Physics, E-ISSN 2399-3650, Vol. 9, no 1, article id 168Article in journal (Refereed) Published
Abstract [en]

As scientific instruments and the literature generate ever larger volumes of data, machine learning (ML) has become essential for organizing, analyzing and interpreting complex information. This Perspective examines how ML accelerates discovery across disciplines, with examples such as brain mapping and exoplanet detection. It also considers situations with different levels of prior knowledge about the underlying phenomenon, outlining strategies to address limitations and exploit ML effectively. Although growing reliance on ML raises challenges for research practice and validation, it is reshaping scientific methods and expanding what can be studied. We also highlight foundation models as a promising route to faster, broader scientific discovery.

Place, publisher, year, edition, pages
Springer Nature, 2026
National Category
Information Systems
Identifiers
urn:nbn:se:kth:diva-385944 (URN)10.1038/s42005-026-02676-7 (DOI)001766603600002 ()2-s2.0-105039471296 (Scopus ID)
Note

QC 20260723

Available from: 2026-07-23 Created: 2026-07-23 Last updated: 2026-07-23Bibliographically approved
Al-Jaff, M., Marchetti, G. L., Welle, M. C., Lundell, J., Gustafsson, M., Henter, G. E., . . . Kragic, D. (2025). A Non-Adversarial Approach to Idempotent Generative Modelling. In: Inês Lynce, Nello Murano, Mauro Vallati, Serena Villata, Federico Chesani, Michela Milano, Andrea Omicini, Mehdi Dastani (Ed.), Proceedings ECAI 2025 - 28th European Conference on Artificial Intelligence: . Paper presented at ECAI 2025 - 28th European Conference on Artificial Intelligence, Including 14th Conference on Prestigious Applications of Intelligent Systems (PAIS 2025), Bologna, Italy, 25-30 October 2025 (pp. 1993-2000). IOS Press, 413, Article ID 10.3233/FAIA251035.
Open this publication in new window or tab >>A Non-Adversarial Approach to Idempotent Generative Modelling
Show others...
2025 (English)In: Proceedings ECAI 2025 - 28th European Conference on Artificial Intelligence / [ed] Inês Lynce, Nello Murano, Mauro Vallati, Serena Villata, Federico Chesani, Michela Milano, Andrea Omicini, Mehdi Dastani, IOS Press , 2025, Vol. 413, p. 1993-2000, article id 10.3233/FAIA251035Conference paper, Published paper (Refereed)
Abstract [en]

Idempotent Generative Networks (IGNs) are deep generative models that also function as local data manifold projectors, mapping arbitrary inputs back onto the manifold. They are trained to act as identity operators on the data and as idempotent operators off the data manifold. However, IGNs suffer from mode collapse, mode dropping, and training instability due to their objectives, which contain adversarial components and can cause the model to cover the data manifold only partially – an issue shared with generative adversarial networks. We introduce Non-Adversarial Idempotent Generative Networks (NAIGNs) to address these issues. Our loss function combines reconstruction with the non-adversarial generative objective of Implicit Maximum Likelihood Estimation (IMLE). This improves on IGN’s ability to restore corrupted data and generate new samples that closely match the data distribution. We moreover demonstrate that NAIGNs implicitly learn the distance field to the data manifold, as well as an energy-based model.

Place, publisher, year, edition, pages
IOS Press, 2025
Series
Frontiers in Artificial Intelligence and Applications, ISSN 0922-6389, E-ISSN 0922-6389 ; 413
Keywords
machine learning, generative modelling
National Category
Electrical Engineering, Electronic Engineering, Information Engineering Computer Vision and Learning Systems
Research subject
Computer Science
Identifiers
urn:nbn:se:kth:diva-384880 (URN)10.3233/FAIA251035 (DOI)001753070100249 ()2-s2.0-105024442848 (Scopus ID)
Conference
ECAI 2025 - 28th European Conference on Artificial Intelligence, Including 14th Conference on Prestigious Applications of Intelligent Systems (PAIS 2025), Bologna, Italy, 25-30 October 2025
Funder
Wallenberg AI, Autonomous Systems and Software Program (WASP)Swedish Research CouncilEU, European Research CouncilKnut and Alice Wallenberg Foundation
Note

Part of ISBN 978-1-64368-631-8

QC 20260706

Available from: 2026-07-06 Created: 2026-07-06 Last updated: 2026-07-27Bibliographically approved
Hafner, S., Fang, H., Azizpour, H. & Ban, Y. (2025). Continuous Urban Change Detection from Satellite Image Time Series with Temporal Feature Refinement and Multi-Task Integration. IEEE Transactions on Geoscience and Remote Sensing, 63, 1-18
Open this publication in new window or tab >>Continuous Urban Change Detection from Satellite Image Time Series with Temporal Feature Refinement and Multi-Task Integration
2025 (English)In: IEEE Transactions on Geoscience and Remote Sensing, ISSN 0196-2892, E-ISSN 1558-0644, Vol. 63, p. 1-18Article in journal (Refereed) Published
Abstract [en]

Urbanization advances at unprecedented rates, leading to negative environmental and societal impacts. Remote sensing can help mitigate these effects by supporting sustainable development strategies with accurate information on urban growth. Deep learning-based methods have achieved promising urban change detection results from optical satellite image pairs using convolutional neural networks (ConvNets), transformers, and a multi-task learning setup. However, bi-temporal methods are limited for continuous urban change detection, i.e., the detection of changes in consecutive image pairs of satellite image time series (SITS), as they fail to fully exploit multi-temporal data (> 2 images). Existing multi-temporal change detection methods, on the other hand, collapse the temporal dimension, restricting their ability to capture continuous urban changes. Additionally, multi-task learning methods lack integration approaches that combine change and segmentation outputs. To address these challenges, we propose a continuous urban change detection framework incorporating two key modules. The temporal feature refinement (TFR) module employs self-attention to improve ConvNet-based multi-temporal building representations. The temporal dimension is preserved in the TFR module, enabling the detection of continuous changes. The multi-task integration (MTI) module utilizes Markov networks to find an optimal building map time series based on segmentation and dense change outputs. The proposed framework effectively identifies urban changes based on high-resolution SITS acquired by the PlanetScope constellation (F1 score 0.551), Gaofen-2 (F1 score 0.440), and WorldView-2 (F1 score 0.543). Moreover, our experiments on three challenging datasets demonstrate the effectiveness of the proposed framework compared to bi-temporal and multi-temporal urban change detection and segmentation methods.

Place, publisher, year, edition, pages
Institute of Electrical and Electronics Engineers (IEEE), 2025
Keywords
Earth observation, Multi-task learning, Multi-temporal, Remote sensing, Transformers
National Category
Earth Observation
Identifiers
urn:nbn:se:kth:diva-366565 (URN)10.1109/TGRS.2025.3578866 (DOI)001512531900009 ()2-s2.0-105007921391 (Scopus ID)
Note

QC 20250710

Available from: 2025-07-10 Created: 2025-07-10 Last updated: 2025-11-03Bibliographically approved
Zhou, W., Sprague, C. I., Viliuga, V., Tadiello, M., Elofsson, A. & Azizpour, H. (2025). Energy-Based Flow Matching for Generating 3D Molecular Structure. In: Proceedings of Machine Learning Research - International Conference on Machine Learning, ICML 2025: . Paper presented at 42nd International Conference on Machine Learning, ICML 2025, Vancouver, Canada, Jul 13 2025 - Jul 19 2025 (pp. 79168-79191). ML Research Press, 267
Open this publication in new window or tab >>Energy-Based Flow Matching for Generating 3D Molecular Structure
Show others...
2025 (English)In: Proceedings of Machine Learning Research - International Conference on Machine Learning, ICML 2025, ML Research Press , 2025, Vol. 267, p. 79168-79191Conference paper, Published paper (Refereed)
Abstract [en]

Molecular structure generation is a fundamental problem that involves determining the 3D positions of molecules’ constituents. It has crucial biological applications, such as molecular docking, protein folding, and molecular design. Recent advances in generative modeling, such as diffusion models and flow matching, have made great progress on these tasks by modeling molecular conformations as a distribution. In this work, we focus on flow matching and adopt an energybased perspective to improve training and inference of structure generation models. Our view results in a mapping function, represented by a deep network, that is directly learned to iteratively map random configurations, i.e. samples from the source distribution, to target structures, i.e. points in the data manifold. This yields a conceptually simple and empirically effective flow matching setup that is theoretically justified and has interesting connections to fundamental properties such as idempotency and stability, as well as the empirically useful techniques such as structure refinement in AlphaFold. Experiments on protein docking as well as protein backbone generation consistently demonstrate the method’s effectiveness, where it outperforms recent baselines of task-associated flow matching and diffusion models, using a similar computational budget.

Place, publisher, year, edition, pages
ML Research Press, 2025
National Category
Computer Sciences
Identifiers
urn:nbn:se:kth:diva-377746 (URN)2-s2.0-105023513160 (Scopus ID)
Conference
42nd International Conference on Machine Learning, ICML 2025, Vancouver, Canada, Jul 13 2025 - Jul 19 2025
Note

QC 20260305

Available from: 2026-03-05 Created: 2026-03-05 Last updated: 2026-03-05Bibliographically approved
Zhou, W., Sprague, C., Viliuga, V., Tadiello, M., Elofsson, A. & Azizpour, H. (2025). Energy-Based Flow Matching for Generating 3D Molecular Structure. In: Singh, A Fazel, M Hsu, D Lacoste-Julien, S Berkenkamp, F Maharaj, T Wagstaff, K Zhu, J (Ed.), International Conference On Machine Learning: . Paper presented at 42nd International Conference on Machine Learning-ICML-Annual, JUL 13-19, 2025, Vancouver, CANADA (pp. 79168-79191). JMLR-JOURNAL MACHINE LEARNING RESEARCH, 267
Open this publication in new window or tab >>Energy-Based Flow Matching for Generating 3D Molecular Structure
Show others...
2025 (English)In: International Conference On Machine Learning / [ed] Singh, A Fazel, M Hsu, D Lacoste-Julien, S Berkenkamp, F Maharaj, T Wagstaff, K Zhu, J, JMLR-JOURNAL MACHINE LEARNING RESEARCH , 2025, Vol. 267, p. 79168-79191Conference paper, Published paper (Refereed)
Abstract [en]

Molecular structure generation is a fundamental problem that involves determining the 3D positions of molecules' constituents. It has crucial biological applications, such as molecular docking, protein folding, and molecular design. Recent advances in generative modeling, such as diffusion models and flow matching, have made great progress on these tasks by modeling molecular conformations as a distribution. In this work, we focus on flow matching and adopt an energy-based perspective to improve training and inference of structure generation models. Our view results in a mapping function, represented by a deep network, that is directly learned to iteratively map random configurations, i.e. samples from the source distribution, to target structures, i.e. points in the data manifold. This yields a conceptually simple and empirically effective flow matching setup that is theoretically justified and has interesting connections to fundamental properties such as idempotency and stability, as well as the empirically useful techniques such as structure refinement in AlphaFold. Experiments on protein docking as well as protein backbone generation consistently demonstrate the method's effectiveness, where it outperforms recent baselines of task-associated flow matching and diffusion models, using a similar computational budget.

Place, publisher, year, edition, pages
JMLR-JOURNAL MACHINE LEARNING RESEARCH, 2025
Series
Proceedings of Machine Learning Research, ISSN 2640-3498
National Category
Computer Sciences
Identifiers
urn:nbn:se:kth:diva-383164 (URN)001693172600173 ()
Conference
42nd International Conference on Machine Learning-ICML-Annual, JUL 13-19, 2025, Vancouver, CANADA
Note

QC 20260610

Available from: 2026-06-10 Created: 2026-06-10 Last updated: 2026-06-10Bibliographically approved
Guastoni, L., Geetha Balasubramanian, A., Foroozan, F., Güemes, A., Ianiro, A., Discetti, S., . . . Vinuesa, R. (2025). Fully convolutional networks for velocity-field predictions based on the wall heat flux in turbulent boundary layers. Theoretical and Computational Fluid Dynamics, 39(1), Article ID 13.
Open this publication in new window or tab >>Fully convolutional networks for velocity-field predictions based on the wall heat flux in turbulent boundary layers
Show others...
2025 (English)In: Theoretical and Computational Fluid Dynamics, ISSN 0935-4964, E-ISSN 1432-2250, Vol. 39, no 1, article id 13Article in journal (Refereed) Published
Abstract [en]

Fully-convolutional neural networks (FCN) were proven to be effective for predicting the instantaneous state of a fully-developed turbulent flow at different wall-normal locations using quantities measured at the wall. In Guastoni et al. (J Fluid Mech 928:A27, 2021. https://doi.org/10.1017/jfm.2021.812), we focused on wall-shear-stress distributions as input, which are difficult to measure in experiments. In order to overcome this limitation, we introduce a model that can take as input the heat-flux field at the wall from a passive scalar. Four different Prandtl numbers Pr=ν/α=(1,2,4,6) are considered (where ν is the kinematic viscosity and α is the thermal diffusivity of the scalar quantity). A turbulent boundary layer is simulated since accurate heat-flux measurements can be performed in experimental settings: first we train the network on aptly-modified DNS data and then we fine-tune it on the experimental data. Finally, we test our network on experimental data sampled in a water tunnel. These predictions represent the first application of transfer learning on experimental data of neural networks trained on simulations. This paves the way for the implementation of a non-intrusive sensing approach for the flow in practical applications.

Place, publisher, year, edition, pages
Springer Nature, 2025
Keywords
Machine learning, Turbulence simulation, Turbulent boundary layers
National Category
Fluid Mechanics
Identifiers
urn:nbn:se:kth:diva-358176 (URN)10.1007/s00162-024-00732-y (DOI)001378464000001 ()2-s2.0-85212435435 (Scopus ID)
Note

Not duplicate with DiVA 1756843

QC 20250114

Available from: 2025-01-07 Created: 2025-01-07 Last updated: 2025-02-09Bibliographically approved
Gutha, S. B., Vinuesa, R. & Azizpour, H. (2025). Inverse Problems with Diffusion Models: A MAP Estimation Perspective. In: Proceedings - 2025 IEEE Winter Conference on Applications of Computer Vision, WACV 2025: . Paper presented at 2025 IEEE/CVF Winter Conference on Applications of Computer Vision, WACV 2025, Tucson, United States of America, Feb 28 2025 - Mar 4 2025 (pp. 4153-4162). Institute of Electrical and Electronics Engineers (IEEE)
Open this publication in new window or tab >>Inverse Problems with Diffusion Models: A MAP Estimation Perspective
2025 (English)In: Proceedings - 2025 IEEE Winter Conference on Applications of Computer Vision, WACV 2025, Institute of Electrical and Electronics Engineers (IEEE) , 2025, p. 4153-4162Conference paper, Published paper (Refereed)
Abstract [en]

Inverse problems have many applications in science and engineering. In Computer vision, several image restoration tasks such as inpainting, deblurring, and super-resolution can be formally modeled as inverse problems. Recently, methods have been developed for solving inverse problems that only leverage a pre-trained unconditional diffusion model and do not require additional task-specific training. In such methods, however, the inherent intractability of determining the conditional score function during the reverse diffusion process poses a real challenge, leaving the methods to settle with an approximation instead, which affects their performance in practice. Here, we propose a MAP estimation framework to model the reverse conditional generation process of a continuous time diffusion model as an optimization process of the underlying MAP objective, whose gradient term is tractable. In theory, the proposed framework can be applied to solve general inverse problems using gradient-based optimization methods. However, given the highly non-convex nature of the loss objective, finding a perfect gradient-based optimization algorithm can be quite challenging, nevertheless, our framework offers several potential research directions. We use our proposed formulation to develop empirically effective algorithms for image restoration. We validate our proposed algorithms with extensive experiments over multiple datasets across several restoration tasks.

Place, publisher, year, edition, pages
Institute of Electrical and Electronics Engineers (IEEE), 2025
Keywords
conditional generation, consistency models, diffusion models, inverse problems, map estimation, optimization
National Category
Computational Mathematics Computer graphics and computer vision
Identifiers
urn:nbn:se:kth:diva-363209 (URN)10.1109/WACV61041.2025.00408 (DOI)001481328900398 ()2-s2.0-105003630084 (Scopus ID)
Conference
2025 IEEE/CVF Winter Conference on Applications of Computer Vision, WACV 2025, Tucson, United States of America, Feb 28 2025 - Mar 4 2025
Note

Part of ISBN 9798331510831

QC 20250509

Available from: 2025-05-07 Created: 2025-05-07 Last updated: 2025-12-08Bibliographically approved
Fang, H. & Azizpour, H. (2025). Leveraging Satellite Image Time Series for Accurate Extreme Event Detection. In: 2025 Ieee/Cvf Winter Conference On Applications Of Computer Vision Workshops, Wacvw: . Paper presented at 2025 Winter Conference on Applications of Computer Vision Workshops-WACVW, FEB 28-MAR 04, 2025, Tucson, AZ (pp. 489-498). Institute of Electrical and Electronics Engineers (IEEE)
Open this publication in new window or tab >>Leveraging Satellite Image Time Series for Accurate Extreme Event Detection
2025 (English)In: 2025 Ieee/Cvf Winter Conference On Applications Of Computer Vision Workshops, Wacvw, Institute of Electrical and Electronics Engineers (IEEE) , 2025, p. 489-498Conference paper, Published paper (Refereed)
Abstract [en]

Climate change is leading to an increase in extreme weather events, causing significant environmental damage and loss of life. Early detection of such events is essential for improving disaster response. In this work, we propose SITS-Extreme, a novel framework that leverages satellite image time series to detect extreme events by incorporating multiple pre-disaster observations. This approach effectively filters out irrelevant changes while isolating disaster-relevant signals, enabling more accurate detection. Extensive experiments on both real-world and synthetic datasets validate the effectiveness of SITS-Extreme, demonstrating substantial improvements over widely used strong bi-temporal baselines. Additionally, we examine the impact of incorporating more timesteps, analyze the contribution of key components in our framework, and evaluate its performance across different disaster types, offering valuable insights into its scalability and applicability for large-scale disaster monitoring.

Place, publisher, year, edition, pages
Institute of Electrical and Electronics Engineers (IEEE), 2025
Series
IEEE Winter Conference on Applications of Computer Vision Workshops, ISSN 2572-4398
National Category
Computer Sciences
Identifiers
urn:nbn:se:kth:diva-373058 (URN)10.1109/WACVW65960.2025.00060 (DOI)001510213100051 ()2-s2.0-105005024910 (Scopus ID)
Conference
2025 Winter Conference on Applications of Computer Vision Workshops-WACVW, FEB 28-MAR 04, 2025, Tucson, AZ
Note

Part of proceedings ISBN 979-833153662-6

QC 20251118

Available from: 2025-11-18 Created: 2025-11-18 Last updated: 2025-11-18Bibliographically approved
Hakkinen, I., Melekhov, I., Englesson, E., Azizpour, H. & Kannala, J. (2025). Medical Image Segmentation with SAM-Generated Annotations. In: DelBue, A Canton, C Pont-Tuset, J Tommasi, T (Ed.), Computer Vision-Eccv 2024 Workshops, Pt Xxii: . Paper presented at 18th European Conference on Computer Vision (ECCV), Sep 29- 04, 2024, Milan, Italy (pp. 51-62). Springer Nature, 15644
Open this publication in new window or tab >>Medical Image Segmentation with SAM-Generated Annotations
Show others...
2025 (English)In: Computer Vision-Eccv 2024 Workshops, Pt Xxii / [ed] DelBue, A Canton, C Pont-Tuset, J Tommasi, T, Springer Nature , 2025, Vol. 15644, p. 51-62Conference paper, Published paper (Refereed)
Abstract [en]

The field of medical image segmentation is hindered by the scarcity of large, publicly available annotated datasets. Not all datasets are made public for privacy reasons, and creating annotations for a large dataset is time-consuming and expensive, as it requires specialized expertise to accurately identify regions of interest (ROIs) within the images. To address these challenges, we evaluate the performance of the Segment Anything Model (SAM) as an annotation tool for medical data by using it to produce so-called "pseudo labels" on the Medical Segmentation Decathlon (MSD) computed tomography (CT) tasks. The pseudo labels are then used in place of ground truth labels to train a UNet model in a weakly-supervised manner. We experiment with different prompt types on SAM and find that the bounding box prompt is a simple yet effective method for generating pseudo labels. This method allows us to develop a weakly-supervised model that performs comparably to a fully supervised model.

Place, publisher, year, edition, pages
Springer Nature, 2025
Series
Lecture Notes in Computer Science, ISSN 0302-9743
Keywords
Foundation Model, Segment Anything Model, Medical Image Segmentation, Data Annotation
National Category
Computer Sciences
Identifiers
urn:nbn:se:kth:diva-374359 (URN)10.1007/978-3-031-92089-9_4 (DOI)001544995100004 ()2-s2.0-105006891572 (Scopus ID)
Conference
18th European Conference on Computer Vision (ECCV), Sep 29- 04, 2024, Milan, Italy
Note

Part of ISBN 978-3-031-92088-2; 978-3-031-92089-9

QC 20251218

Available from: 2025-12-18 Created: 2025-12-18 Last updated: 2025-12-18Bibliographically approved
Mehrpanah, A., Englesson, E. & Azizpour, H. (2025). On Spectral Properties of Gradient-Based Explanation Methods. In: Computer Vision – ECCV 2024 - 18th European Conference, Proceedings: . Paper presented at 18th European Conference on Computer Vision, ECCV 2024, Milan, Italy, Sep 29 2024 - Oct 4 2024 (pp. 282-299). Springer Nature
Open this publication in new window or tab >>On Spectral Properties of Gradient-Based Explanation Methods
2025 (English)In: Computer Vision – ECCV 2024 - 18th European Conference, Proceedings, Springer Nature , 2025, p. 282-299Conference paper, Published paper (Refereed)
Abstract [en]

Understanding the behavior of deep networks is crucial to increase our confidence in their results. Despite an extensive body of work for explaining their predictions, researchers have faced reliability issues, which can be attributed to insufficient formalism. In our research, we adopt novel probabilistic and spectral perspectives to formally analyze explanation methods. Our study reveals a pervasive spectral bias stemming from the use of gradient, and sheds light on some common design choices that have been discovered experimentally, in particular, the use of squared gradient and input perturbation. We further characterize how the choice of perturbation hyperparameters in explanation methods, such as SmoothGrad, can lead to inconsistent explanations and introduce two remedies based on our proposed formalism: (i) a mechanism to determine a standard perturbation scale, and (ii) an aggregation method which we call SpectralLens. Finally, we substantiate our theoretical results through quantitative evaluations.

Place, publisher, year, edition, pages
Springer Nature, 2025
Keywords
Deep Neural Networks, Explainability, Gradient-based Explanation Methods, Probabilistic Machine Learning, Probabilistic Pixel Attribution Techniques, Spectral Analysis
National Category
Computer Sciences
Identifiers
urn:nbn:se:kth:diva-357698 (URN)10.1007/978-3-031-73021-4_17 (DOI)001416940200017 ()2-s2.0-85210488897 (Scopus ID)
Conference
18th European Conference on Computer Vision, ECCV 2024, Milan, Italy, Sep 29 2024 - Oct 4 2024
Note

Part of ISBN 978-303173020-7

QC 20241213

Available from: 2024-12-12 Created: 2024-12-12 Last updated: 2025-03-17Bibliographically approved
Organisations
Identifiers
ORCID iD: ORCID iD iconorcid.org/0000-0001-5211-6388

Search in DiVA

Show all publications