kth.sePublications KTH
Change search
Link to record
Permanent link

Direct link
Publications (10 of 62) Show all publications
Szczerek, W. J. & Podobas, A. (2026). A Quarter of a Century of Neuromorphic Architectures on FPGAs - An Overview. ACM Computing Surveys, Article ID 3831666.
Open this publication in new window or tab >>A Quarter of a Century of Neuromorphic Architectures on FPGAs - An Overview
2026 (English)In: ACM Computing Surveys, ISSN 0360-0300, E-ISSN 1557-7341, article id 3831666Article, review/survey (Refereed) Epub ahead of print
Abstract [en]

Neuromorphic computing is a relatively new discipline of computer science, where the principles of biological brain’s computation and memory are used to create a new way of processing information, based on networks of spiking neurons. Those networks can be implemented as both analog and digital implementations, where for the latter, the Field Programmable Gate Arrays (FPGAs) are a frequent choice, due to their inherent flexibility, allowing the researchers to easily design hardware neuromorphic architecture (NMAs). Moreover, digital NMAs show good promise in simulating various spiking neural networks because of their inherent accuracy and resilience to noise, as opposed to analog implementations. This paper presents an overview of digital NMAs implemented on FPGAs, with a goal of providing useful references to various architectural design choices to the researchers interested in digital neuromorphic systems. We present a taxonomy of NMAs that highlights groups of distinct architectural features, their advantages and disadvantages and identify trends and predictions for the future of those architectures. 

Place, publisher, year, edition, pages
Association for Computing Machinery (ACM), 2026
National Category
Computer Systems
Research subject
Computer Science; Information and Communication Technology
Identifiers
urn:nbn:se:kth:diva-385569 (URN)10.1145/3831666 (DOI)
Projects
Building Digital Brains
Funder
Swedish Research Council, 2021-04579
Note

QC 20260715

Available from: 2026-07-15 Created: 2026-07-15 Last updated: 2026-07-15Bibliographically approved
Gonidakis, P., Carella, F., Dineva, E., Jeong, H. J., Antunes, P., Podobas, A., . . . Miloshevich, G. (2026). Comparing Solar Structure Detection Methods in SDO/AIA Observations and the Application to Raw Uncalibrated Data. Journal of Geophysical Research: Machine Learning and Computation:, 3(3), Article ID e2025JH001023.
Open this publication in new window or tab >>Comparing Solar Structure Detection Methods in SDO/AIA Observations and the Application to Raw Uncalibrated Data
Show others...
2026 (English)In: Journal of Geophysical Research: Machine Learning and Computation:, E-ISSN 2993-5210, Vol. 3, no 3, article id e2025JH001023Article in journal (Refereed) Published
Abstract [en]

Recent advances in solar physics increasingly rely on automated identification of coronal structures using machine learning. Yet most studies emphasize scientific performance without evaluating feasibility for onboard deployment to prioritize downlink observations. This work investigates the automated identification of active regions and coronal holes by applying segmentation and detection techniques to Solar Dynamics Observatory (SDO) data. We compare three approaches: SCSS-Net, a deep learning model for semantic segmentation; YOLOv8n, a lightweight object detector; and a traditional pipeline based on basic computer vision operations (BCVO). Each method is assessed for its scientific accuracy and its suitability for deployment in future resource-limited missions. While no direct hardware benchmarking is performed in this study, we assess the feasibility of onboard implementation based on the associated number of trainable parameters, architecture and hardware requirements. Training and evaluation are first conducted on well-calibrated SDO images. We then extend the evaluation to raw and uncalibrated SDO images affected by instrumental artifacts. Performance is measured using the Intersection over Union (IoU) and Dice score. Results show that while SCSS-Net achieves the highest segmentation quality, YOLOv8n offers a strong balance between accuracy and efficiency. The BCVO pipeline remains viable under strict hardware limitations. Interestingly, our models retain compatibility on Level-0 observations. This is the first study comparing these widely used methods from the perspective of onboard deployment. Our findings provide a foundation for designing frameworks tailored to onboard hardware configurations.

Place, publisher, year, edition, pages
American Geophysical Union (AGU), 2026
Keywords
Level-0 SDO, SDO, active regions, coronal holes, machine learning, solar coronal structure detection
National Category
Computer graphics and computer vision Astronomy, Astrophysics and Cosmology
Identifiers
urn:nbn:se:kth:diva-383049 (URN)10.1029/2025JH001023 (DOI)2-s2.0-105039875481 (Scopus ID)
Note

QC 20260605

Available from: 2026-06-05 Created: 2026-06-05 Last updated: 2026-06-05Bibliographically approved
Chien, S. W. .., Sato, K., Podobas, A., Jansson, N., Markidis, S. & Honda, M. (2026). ParaLog: Consistent Host-side Logging for Parallel Checkpoints. In: SoCC 2025 - Proceedings of the 2025 ACM Symposium on Cloud Computing: . Paper presented at 2025 ACM Symposium on Cloud Computing, SoCC 2025, Virtual, Online, United States of America, November 19-21, 2025 (pp. 59-73). Association for Computing Machinery, Inc
Open this publication in new window or tab >>ParaLog: Consistent Host-side Logging for Parallel Checkpoints
Show others...
2026 (English)In: SoCC 2025 - Proceedings of the 2025 ACM Symposium on Cloud Computing, Association for Computing Machinery, Inc , 2026, p. 59-73Conference paper, Published paper (Refereed)
Abstract [en]

Output-intensive scientific applications are highly sensitive to low storage throughput. While existing scientific application stacks are optimized for traditional High-Performance Computing (HPC) environments with high remote storage and network bandwidth, these assumptions often fail in modern settings like cloud deployment. This is because the existing scientific application I/O stack fails to leverage the available resources. At the same time, scientific applications exhibit special synchronization and data output requirements that are difficult to satisfy using traditional approaches such as block-level or filesystem-level caching. We introduce ParaLog, a distributed host-side logging approach designed to accelerate scientific applications transparently. ParaLog emphasizes deployability, enabling support for unmodified message passing interface (MPI) applications and implementations while preserving crash consistency semantics. We evaluate ParaLog across traditional HPC, cloud HPC, local clusters, and hybrid environments, demonstrating its capability to reduce end-to-end execution time by 13-26% for popular scientific applications in cloud settings.

Place, publisher, year, edition, pages
Association for Computing Machinery, Inc, 2026
Keywords
burst buffer, caching, Cloud Computing, High Performance Computing, parallel IO, S3, scientific applications
National Category
Computer Sciences Computer Systems
Identifiers
urn:nbn:se:kth:diva-376725 (URN)10.1145/3772052.3772212 (DOI)001697656400005 ()2-s2.0-105028598983 (Scopus ID)
Conference
2025 ACM Symposium on Cloud Computing, SoCC 2025, Virtual, Online, United States of America, November 19-21, 2025
Note

Part of ISBN 9798400722769

QC 20260213

Available from: 2026-02-13 Created: 2026-02-13 Last updated: 2026-05-29Bibliographically approved
Al Hafiz, M. I., Ravichandran, N. B., Lansner, A., Herman, P. & Podobas, A. (2025). A Reconfigurable Stream-Based FPGA Accelerator for Bayesian Confidence Propagation Neural Networks. In: Applied Reconfigurable Computing. Architectures, Tools, and Applications - 21st International Symposium, ARC 2025, Proceedings: . Paper presented at 21st International Symposium on Applied Reconfigurable Computing, ARC 2025, Seville, Spain, Apr 9 2025 - Apr 11 2025 (pp. 196-213). Springer Nature
Open this publication in new window or tab >>A Reconfigurable Stream-Based FPGA Accelerator for Bayesian Confidence Propagation Neural Networks
Show others...
2025 (English)In: Applied Reconfigurable Computing. Architectures, Tools, and Applications - 21st International Symposium, ARC 2025, Proceedings, Springer Nature , 2025, p. 196-213Conference paper, Published paper (Refereed)
Abstract [en]

Brain-like algorithms are attractive and emerging alternatives to classical deep learning methods for use in various machine learning applications. Brain-like systems can feature local learning rules, both unsupervised/semi-supervised learning and different types of plasticity (structural/synaptic), allowing them to potentially be faster and more energy-efficient than traditional machine learning alternatives. Among the more salient brain-like algorithms are Bayesian Confidence Propagation Neural Networks (BCPNNs). BCPNN is an important tool for both machine learning and computational neuroscience research, and recent work shows that BCPNN can reach state-of-the-art performance in tasks such as learning and memory recall compared to other models. Unfortunately, BCPNN is primarily executed on slow general-purpose processors (CPUs) or power-hungry graphics processing units (GPUs), reducing the applicability of using BCPNN in Edge systems, among others. In this work, we design a reconfigurable stream-based accelerator for BCPNN using Field-Programmable Gate Arrays (FPGA) using Xilinx Vitis High-Level Synthesis (HLS) flow. Furthermore, we model our accelerator’s performance using first principles, and we empirically show that our proposed accelerator (full-featured kernel non-structural plasticity) is between 1.3x - 5.3x faster than an Nvidia A100 GPU while at the same time consuming between 2.62x - 3.19x less power and 5.8x - 16.5x less energy without any degradation in performance.

Place, publisher, year, edition, pages
Springer Nature, 2025
Keywords
BCPNN, FPGA, HLS, Neuromorphic
National Category
Computer Sciences Other Electrical Engineering, Electronic Engineering, Information Engineering
Identifiers
urn:nbn:se:kth:diva-363095 (URN)10.1007/978-3-031-87995-1_12 (DOI)2-s2.0-105002874652 (Scopus ID)
Conference
21st International Symposium on Applied Reconfigurable Computing, ARC 2025, Seville, Spain, Apr 9 2025 - Apr 11 2025
Note

Part of ISBN 9783031879944

QC 20250922

Available from: 2025-05-06 Created: 2025-05-06 Last updated: 2025-09-22Bibliographically approved
Al Hafiz, M. I., Ravichandran, N. B., Lansner, A., Herman, P. & Podobas, A. (2025). Embedded FPGA Acceleration of Brain-Like Neural Networks: Online Learning to Scalable Inference. In: Proc. - IEEE Int. Symp. Embed. Multicore/Many-core Syst.-on-Chip, MCSoC: . Paper presented at 18th IEEE International Symposium on Embedded Multicore/Many-core Systems-on-Chip, MCSoC 2025, Singapore, December 15-18, 2025 (pp. 331-338). Institute of Electrical and Electronics Engineers (IEEE)
Open this publication in new window or tab >>Embedded FPGA Acceleration of Brain-Like Neural Networks: Online Learning to Scalable Inference
Show others...
2025 (English)In: Proc. - IEEE Int. Symp. Embed. Multicore/Many-core Syst.-on-Chip, MCSoC, Institute of Electrical and Electronics Engineers (IEEE) , 2025, p. 331-338Conference paper, Published paper (Refereed)
Abstract [en]

Edge AI increasingly requires models that learn and adapt on-device under a tight energy budget. Mainstream deep learning models, while powerful, are often overparameterized, energy-hungry and dependent on cloud connectivity. Brain-Like Neural Networks (BLNNs), such as the Bayesian Confidence Propagation Neural Network (BCPNN), propose a neuromorphic alternative by mimicking cortical architecture and biologicallyconstrained learning. They offer sparse architectures with local learning rules and unsupervised/semi-supervised learning, making them well-suited for low-power edge intelligence. However, existing BCPNN implementations rely on GPUs or datacenter FPGAs. This work presents the first embedded FPGA accelerator for BCPNN on a Zynq UltraScale+ SoC (ZCU104) using High-Level Synthesis. We implement both online learning and inference-only kernels with configurable precision (FP32, FP16, and mixed FP16/FXP16). Evaluated on MNIST, Pneumonia, and Breast Cancer datasets, our accelerator delivers up to 17.55% lower latency and 94.1% energy savings over ARM baselines. Our work brings practical, brain-like online learning and scalable inference to edge devices.

Place, publisher, year, edition, pages
Institute of Electrical and Electronics Engineers (IEEE), 2025
Keywords
BCPNN, BLNN, Embedded, FPGA, HLS, Neuromorphic, Bayesian networks, Brain, Budget control, Deep learning, E-learning, High level synthesis, Logic Synthesis, Low power electronics, Network architecture, Neural networks, Online systems, Program processors, Programmable logic controllers, System-on-chip, Bayesian, Bayesian confidence propagation neural network, Brain-like neural network, Embedded FPGA, Neural-networks, Online learning, Field programmable gate arrays (FPGA)
National Category
Computer Sciences
Identifiers
urn:nbn:se:kth:diva-384584 (URN)10.1109/MCSoC67473.2025.00060 (DOI)2-s2.0-105032396875 (Scopus ID)
Conference
18th IEEE International Symposium on Embedded Multicore/Many-core Systems-on-Chip, MCSoC 2025, Singapore, December 15-18, 2025
Note

Part of ISBN 9798331565718

QC 20260716

Available from: 2026-07-01 Created: 2026-07-01 Last updated: 2026-07-16Bibliographically approved
Antunes, P., Al Hafiz, M. I., Ekelund, J., Dineva, E., Miloshevich, G., Gonidakis, P. & Podobas, A. (2025). Evaluating Four FPGA-Accelerated Space Use Cases Based on Neural Network Algorithms for On-Board Inference. In: Proc. - IEEE Int. Symp. Embed. Multicore/Many-core Syst.-on-Chip, MCSoC: . Paper presented at 18th IEEE International Symposium on Embedded Multicore/Many-core Systems-on-Chip, MCSoC 2025, Singapore, December 15-18, 2025 (pp. 804-812). Institute of Electrical and Electronics Engineers (IEEE)
Open this publication in new window or tab >>Evaluating Four FPGA-Accelerated Space Use Cases Based on Neural Network Algorithms for On-Board Inference
Show others...
2025 (English)In: Proc. - IEEE Int. Symp. Embed. Multicore/Many-core Syst.-on-Chip, MCSoC, Institute of Electrical and Electronics Engineers (IEEE) , 2025, p. 804-812Conference paper, Published paper (Refereed)
Abstract [en]

Space missions increasingly deploy high-fidelity sensors that produce data volumes exceeding onboard buffering and downlink capacity. This work evaluates FPGA acceleration of neural networks (NNs) across four space use cases on the AMD ZCU104 board. We use Vitis AI (AMD DPU) and Vitis HLS to implement inference, quantify throughput and energy, and expose toolchain and architectural constraints relevant to deployment. Vitis AI achieves up to 34.16 × higher inference rate than the embedded ARM CPU baseline, while custom HLS designs reach up to 5.4 × speedup and add support for operators (e.g., sigmoids, 3D layers) absent in the DPU. For these implementations, measured MPSoC inference power spans 1.56.75 W, reducing energy per inference versus CPU execution in all use cases. These results show that NN FPGA acceleration can enable onboard filtering, compression, and event detection, easing downlink pressure in future missions. 

Place, publisher, year, edition, pages
Institute of Electrical and Electronics Engineers (IEEE), 2025
Keywords
FPGA, HLS, Neural Network, Space Mission, Vitis AI, Embedded systems, Inference engines, Neural networks, Signal processing, Space flight, Case based, Data volume, Energy, High-fidelity, Neural networks algorithms, Neural-networks, Space missions, Space use, Field programmable gate arrays (FPGA)
National Category
Computer Sciences
Identifiers
urn:nbn:se:kth:diva-384585 (URN)10.1109/MCSoC67473.2025.00126 (DOI)2-s2.0-105032402760 (Scopus ID)
Conference
18th IEEE International Symposium on Embedded Multicore/Many-core Systems-on-Chip, MCSoC 2025, Singapore, December 15-18, 2025
Note

Part of ISBN 9798331565718

QC 20260709

Available from: 2026-07-01 Created: 2026-07-01 Last updated: 2026-07-16Bibliographically approved
Lindqvist, B. A. & Podobas, A. (2025). IncineRate: Multi-Modal FPGA Accelerator Architecture for SCNNs. In: Proceedings - 2025 IEEE 18th International Symposium on Embedded Multicore/Many-core Systems-on-Chip, MCSoC 2025: . Paper presented at 18th International Symposium on Embedded Multicore/Many-core Systems-on-Chip, MCSoC 2025, Singapore, Singapore, December 15-18, 2025 (pp. 165-172). Institute of Electrical and Electronics Engineers (IEEE)
Open this publication in new window or tab >>IncineRate: Multi-Modal FPGA Accelerator Architecture for SCNNs
2025 (English)In: Proceedings - 2025 IEEE 18th International Symposium on Embedded Multicore/Many-core Systems-on-Chip, MCSoC 2025, Institute of Electrical and Electronics Engineers (IEEE) , 2025, p. 165-172Conference paper, Published paper (Refereed)
Abstract [en]

Spiking Neural Networks (SNNs) are a promising alternative to conventional Artificial Neural Networks (ANNs) due to their biological interpretability and capability to exploit sparse computation. Specialized hardware for SNNs has potential advantages over general-purpose devices in terms of power and performance. However, the computational requirements of modern Spiking Convolutional Neural Networks (SCNNs) renders most SNN hardware inefficient for SCNN acceleration. As a step towards efficient SCNN acceleration, we present IncineRate, a flexible FPGA-based SCNN accelerator architecture. IncineRate has built-in support for many SCNNs, such as AlexNet, VGG16, and ResNets, and can be extended to support other network models. The number of simulation time steps, the network architecture, and other settings are specified at run time, so an already deployed device can execute multiple networks without reconfiguration. Our results show that IncineRate achieves state-of-the-art classification accuracy among FPGA-based SCNNs on CIFAR10 and CIFAR100.

Place, publisher, year, edition, pages
Institute of Electrical and Electronics Engineers (IEEE), 2025
Keywords
computer vision, fpga, neuromorphic architecture, reconfigurable hardware, scnn
National Category
Computer Engineering
Identifiers
urn:nbn:se:kth:diva-378751 (URN)10.1109/MCSoC67473.2025.00035 (DOI)2-s2.0-105032437580 (Scopus ID)
Conference
18th International Symposium on Embedded Multicore/Many-core Systems-on-Chip, MCSoC 2025, Singapore, Singapore, December 15-18, 2025
Note

Part of ISBN 9798331565718

QC 20260331

Available from: 2026-03-31 Created: 2026-03-31 Last updated: 2026-03-31Bibliographically approved
Szczerek, W. J. & Podobas, A. (2025). IzhiRISC-V-a RISC-V-based Processor with Custom ISA Extension for Spiking Neuron Networks Processing with Izhikevich Neurons. In: Proceedings of the SC’25 Workshops of the International Conference for High Performance Computing, Networking, Storage and Analysis: . Paper presented at SC Workshops '25: Workshops of the International Conference for High Performance Computing, Networking, Storage and Analysis, St Louis, MO, USA, November 16-21, 2025 (pp. 1667-1675). New York, NY, United States of America: Association for Computing Machinery (ACM)
Open this publication in new window or tab >>IzhiRISC-V-a RISC-V-based Processor with Custom ISA Extension for Spiking Neuron Networks Processing with Izhikevich Neurons
2025 (English)In: Proceedings of the SC’25 Workshops of the International Conference for High Performance Computing, Networking, Storage and Analysis, New York, NY, United States of America: Association for Computing Machinery (ACM) , 2025, p. 1667-1675Conference paper, Published paper (Refereed)
Abstract [en]

Spiking Neural Network processing promises to provide high energy efficiency due to the sparsity of the spiking events. However, when realized on general-purpose hardware – such as a RISC-V processor – this promise can be undermined and overshadowed by the inefficient code, stemming from repeated usage of basic instructions for updating all the neurons in the network. One of the possible solutions to this issue is the introduction of a custom ISA extension with neuromorphic instructions for spiking neuron updating, and realizing those instructions in bespoke hardware expansion to the existing ALU. In this paper, we present the first step towards realizing a large-scale system based on the RISC-V-compliant processor called IzhiRISC-V, supporting the custom neuromorphic ISA extension.

Place, publisher, year, edition, pages
New York, NY, United States of America: Association for Computing Machinery (ACM), 2025
Keywords
RISC-V, ISA, spiking neural networks, neuromorphic computing
National Category
Computer Systems
Identifiers
urn:nbn:se:kth:diva-373252 (URN)10.1145/3731599.3767529 (DOI)001661298800179 ()2-s2.0-105023445932 (Scopus ID)
Conference
SC Workshops '25: Workshops of the International Conference for High Performance Computing, Networking, Storage and Analysis, St Louis, MO, USA, November 16-21, 2025
Funder
Swedish Research Council, 2021-04579
Note

QC 20251210

Available from: 2025-11-25 Created: 2025-11-25 Last updated: 2026-06-22Bibliographically approved
Podobas, A., Sano, K., Anderson, J., Ueno, T., Koch, A., Adhi, B. A., . . . Kojima, T. (2025). The Fourth International Workshop on Coarse-Grained Reconfigurable Architectures for High-Performance Computing and AI (CGRA4HPCA). In: Proceedings 2025 IEEE International Parallel and Distributed Processing Symposium Workshops Ipdpsw 2025: . Paper presented at CGRA4HPCA 2025, Milano, Italy, June 3-7, 2025 (pp. 87-88). Institute of Electrical and Electronics Engineers (IEEE)
Open this publication in new window or tab >>The Fourth International Workshop on Coarse-Grained Reconfigurable Architectures for High-Performance Computing and AI (CGRA4HPCA)
Show others...
2025 (English)In: Proceedings 2025 IEEE International Parallel and Distributed Processing Symposium Workshops Ipdpsw 2025, Institute of Electrical and Electronics Engineers (IEEE) , 2025, p. 87-88Conference paper, Published paper (Refereed)
Place, publisher, year, edition, pages
Institute of Electrical and Electronics Engineers (IEEE), 2025
National Category
Computer Engineering
Identifiers
urn:nbn:se:kth:diva-371026 (URN)10.1109/IPDPSW66978.2025.00019 (DOI)2-s2.0-105015509147 (Scopus ID)
Conference
CGRA4HPCA 2025, Milano, Italy, June 3-7, 2025
Note

Part of ISBN 9798331526436

QC 20251003

Available from: 2025-10-03 Created: 2025-10-03 Last updated: 2025-10-03Bibliographically approved
Lindqvist, B. & Podobas, A. (2024). Algorithms for Fast Spiking Neural Network Simulation on FPGAs. IEEE Access, 12, 150334-150353
Open this publication in new window or tab >>Algorithms for Fast Spiking Neural Network Simulation on FPGAs
2024 (English)In: IEEE Access, E-ISSN 2169-3536, Vol. 12, p. 150334-150353Article in journal (Refereed) Published
Abstract [en]

Spiking Neural Networks (SNNs) are models that mimic and replicate the computational properties of the biological brain. Computation is performed using neurons that transmit information on axons between each other via synapses. SNNs have several important application areas, ranging from (brain-like) artificial intelligence to complex brain simulations. Most SNN simulations today are carried out on systems such as CPUs and GPUs, which fit SNNs poorly and often yield slow solutions that consume needlessly much energy. In this work, we present algorithms for efficient simulation of SNNs on Field-Programmable Gate Arrays (FPGAs), which is driven by our hypothesis that said devices can be much more power-efficient without sacrificing execution performance. We also provide an in-depth analysis and discussion of our algorithms and techniques. We target the important Potjans-Diesmann model, a well-known cortical microcircuit often used for assessing SNN simulation performance. Using high-level synthesis (HLS) targeting the latest Intel Agilex 7 FPGA, we show that our best simulator can execute the microcircuit 25% faster than real-time and require only 21 nJ per synaptic event. Our result surpasses the state-of-the-art for single-device simulation, and the energy use is the lowest among published results.

Place, publisher, year, edition, pages
Institute of Electrical and Electronics Engineers (IEEE), 2024
Keywords
Field programmable gate arrays, Spiking neural networks, Hardware, Synapses, Logic, Membrane potentials, Brain modeling, Table lookup, Neuroscience, Cortical microcircuit, FPGA, HLS, HPC, OpenCL, simulation, leaky integrate-and-fire
National Category
Bioinformatics and Computational Biology
Identifiers
urn:nbn:se:kth:diva-355765 (URN)10.1109/ACCESS.2024.3479933 (DOI)001340664700001 ()2-s2.0-85207714571 (Scopus ID)
Note

QC 20241104

Available from: 2024-11-04 Created: 2024-11-04 Last updated: 2025-05-27Bibliographically approved
Organisations
Identifiers
ORCID iD: ORCID iD iconorcid.org/0000-0001-5452-6794

Search in DiVA

Show all publications