kth.sePublications KTH
Change search
Link to record
Permanent link

Direct link
Al Hafiz, Muhammad IhsanORCID iD iconorcid.org/0000-0002-9150-3847
Publications (4 of 4) Show all publications
Al Hafiz, M. I. & Podobas, A. (2027). NeuroRing: Scaling Spiking Neural Networks via Multi-FPGA Bidirectional Ring Topologies and Stream-Dataflow Architectures. In: Euro-Par 2026: Parallel Processing - 32nd European Conference on Parallel and Distributed Processing, Proceedings. Paper presented at 32nd European Conference on Parallel and Distributed Processing, Euro-Par 2026, Pisa, Italy, August 24-28, 2026 (pp. 329-343). Springer Nature, 16781 LNCS
Open this publication in new window or tab >>NeuroRing: Scaling Spiking Neural Networks via Multi-FPGA Bidirectional Ring Topologies and Stream-Dataflow Architectures
2027 (English)In: Euro-Par 2026: Parallel Processing - 32nd European Conference on Parallel and Distributed Processing, Proceedings, Springer Nature , 2027, Vol. 16781 LNCS, p. 329-343Conference paper, Published paper (Refereed)
Abstract [en]

Spiking neural networks (SNNs) are a promising paradigm for energy-efficient event-driven computation, but large-scale SNN execution remains challenging because sparse spike communication and synchronization can dominate runtime. Existing solutions across CPU, GPU, ASIC, and FPGA platforms offer different trade-offs between programmability, efficiency, and scalability. To address this gap, we present NeuroRing, a modular and scalable SNN accelerator based on a stream-dataflow architecture and a bidirectional ring topology, implemented in High-Level Synthesis (HLS) on FPGAs. NeuroRing supports modular single- and multi-FPGA deployment and is compatible with existing SNN workflows through integration with the NEST simulator. We evaluate NeuroRing on the cortical microcircuit benchmark and a Sudoku constraint-satisfaction workload. Results show that NeuroRing preserves the key activity statistics of the NEST reference model, achieves faster-than-real-time execution of the full-scale cortical microcircuit with a real-time factor (RTF) of 0.83, exhibits meaningful strong and weak scaling, and provides competitive energy efficiency on two programmable FPGAs. These results position NeuroRing as a flexible and scalable platform for both neuroscience simulation and broader event-driven applications.

Place, publisher, year, edition, pages
Springer Nature, 2027
Keywords
Accelerator, Dataflow, FPGA, Ring Topology, SNN, Stream
National Category
Subatomic Physics Computer Sciences Embedded Systems
Identifiers
urn:nbn:se:kth:diva-388190 (URN)10.1007/978-3-032-35248-4_23 (DOI)2-s2.0-105048945351 (Scopus ID)
Conference
32nd European Conference on Parallel and Distributed Processing, Euro-Par 2026, Pisa, Italy, August 24-28, 2026
Note

Part of ISBN 9783032352477

QC 20260911

Available from: 2026-09-11 Created: 2026-09-11 Last updated: 2026-09-11Bibliographically approved
Al Hafiz, M. I., Ravichandran, N. B., Lansner, A., Herman, P. & Podobas, A. (2025). A Reconfigurable Stream-Based FPGA Accelerator for Bayesian Confidence Propagation Neural Networks. In: Applied Reconfigurable Computing. Architectures, Tools, and Applications - 21st International Symposium, ARC 2025, Proceedings: . Paper presented at 21st International Symposium on Applied Reconfigurable Computing, ARC 2025, Seville, Spain, Apr 9 2025 - Apr 11 2025 (pp. 196-213). Springer Nature
Open this publication in new window or tab >>A Reconfigurable Stream-Based FPGA Accelerator for Bayesian Confidence Propagation Neural Networks
Show others...
2025 (English)In: Applied Reconfigurable Computing. Architectures, Tools, and Applications - 21st International Symposium, ARC 2025, Proceedings, Springer Nature , 2025, p. 196-213Conference paper, Published paper (Refereed)
Abstract [en]

Brain-like algorithms are attractive and emerging alternatives to classical deep learning methods for use in various machine learning applications. Brain-like systems can feature local learning rules, both unsupervised/semi-supervised learning and different types of plasticity (structural/synaptic), allowing them to potentially be faster and more energy-efficient than traditional machine learning alternatives. Among the more salient brain-like algorithms are Bayesian Confidence Propagation Neural Networks (BCPNNs). BCPNN is an important tool for both machine learning and computational neuroscience research, and recent work shows that BCPNN can reach state-of-the-art performance in tasks such as learning and memory recall compared to other models. Unfortunately, BCPNN is primarily executed on slow general-purpose processors (CPUs) or power-hungry graphics processing units (GPUs), reducing the applicability of using BCPNN in Edge systems, among others. In this work, we design a reconfigurable stream-based accelerator for BCPNN using Field-Programmable Gate Arrays (FPGA) using Xilinx Vitis High-Level Synthesis (HLS) flow. Furthermore, we model our accelerator’s performance using first principles, and we empirically show that our proposed accelerator (full-featured kernel non-structural plasticity) is between 1.3x - 5.3x faster than an Nvidia A100 GPU while at the same time consuming between 2.62x - 3.19x less power and 5.8x - 16.5x less energy without any degradation in performance.

Place, publisher, year, edition, pages
Springer Nature, 2025
Keywords
BCPNN, FPGA, HLS, Neuromorphic
National Category
Computer Sciences Other Electrical Engineering, Electronic Engineering, Information Engineering
Identifiers
urn:nbn:se:kth:diva-363095 (URN)10.1007/978-3-031-87995-1_12 (DOI)2-s2.0-105002874652 (Scopus ID)
Conference
21st International Symposium on Applied Reconfigurable Computing, ARC 2025, Seville, Spain, Apr 9 2025 - Apr 11 2025
Note

Part of ISBN 9783031879944

QC 20250922

Available from: 2025-05-06 Created: 2025-05-06 Last updated: 2025-09-22Bibliographically approved
Al Hafiz, M. I., Ravichandran, N. B., Lansner, A., Herman, P. & Podobas, A. (2025). Embedded FPGA Acceleration of Brain-Like Neural Networks: Online Learning to Scalable Inference. In: Proc. - IEEE Int. Symp. Embed. Multicore/Many-core Syst.-on-Chip, MCSoC: . Paper presented at 18th IEEE International Symposium on Embedded Multicore/Many-core Systems-on-Chip, MCSoC 2025, Singapore, December 15-18, 2025 (pp. 331-338). Institute of Electrical and Electronics Engineers (IEEE)
Open this publication in new window or tab >>Embedded FPGA Acceleration of Brain-Like Neural Networks: Online Learning to Scalable Inference
Show others...
2025 (English)In: Proc. - IEEE Int. Symp. Embed. Multicore/Many-core Syst.-on-Chip, MCSoC, Institute of Electrical and Electronics Engineers (IEEE) , 2025, p. 331-338Conference paper, Published paper (Refereed)
Abstract [en]

Edge AI increasingly requires models that learn and adapt on-device under a tight energy budget. Mainstream deep learning models, while powerful, are often overparameterized, energy-hungry and dependent on cloud connectivity. Brain-Like Neural Networks (BLNNs), such as the Bayesian Confidence Propagation Neural Network (BCPNN), propose a neuromorphic alternative by mimicking cortical architecture and biologicallyconstrained learning. They offer sparse architectures with local learning rules and unsupervised/semi-supervised learning, making them well-suited for low-power edge intelligence. However, existing BCPNN implementations rely on GPUs or datacenter FPGAs. This work presents the first embedded FPGA accelerator for BCPNN on a Zynq UltraScale+ SoC (ZCU104) using High-Level Synthesis. We implement both online learning and inference-only kernels with configurable precision (FP32, FP16, and mixed FP16/FXP16). Evaluated on MNIST, Pneumonia, and Breast Cancer datasets, our accelerator delivers up to 17.55% lower latency and 94.1% energy savings over ARM baselines. Our work brings practical, brain-like online learning and scalable inference to edge devices.

Place, publisher, year, edition, pages
Institute of Electrical and Electronics Engineers (IEEE), 2025
Keywords
BCPNN, BLNN, Embedded, FPGA, HLS, Neuromorphic, Bayesian networks, Brain, Budget control, Deep learning, E-learning, High level synthesis, Logic Synthesis, Low power electronics, Network architecture, Neural networks, Online systems, Program processors, Programmable logic controllers, System-on-chip, Bayesian, Bayesian confidence propagation neural network, Brain-like neural network, Embedded FPGA, Neural-networks, Online learning, Field programmable gate arrays (FPGA)
National Category
Computer Sciences
Identifiers
urn:nbn:se:kth:diva-384584 (URN)10.1109/MCSoC67473.2025.00060 (DOI)2-s2.0-105032396875 (Scopus ID)
Conference
18th IEEE International Symposium on Embedded Multicore/Many-core Systems-on-Chip, MCSoC 2025, Singapore, December 15-18, 2025
Note

Part of ISBN 9798331565718

QC 20260716

Available from: 2026-07-01 Created: 2026-07-01 Last updated: 2026-07-16Bibliographically approved
Antunes, P., Al Hafiz, M. I., Ekelund, J., Dineva, E., Miloshevich, G., Gonidakis, P. & Podobas, A. (2025). Evaluating Four FPGA-Accelerated Space Use Cases Based on Neural Network Algorithms for On-Board Inference. In: Proc. - IEEE Int. Symp. Embed. Multicore/Many-core Syst.-on-Chip, MCSoC: . Paper presented at 18th IEEE International Symposium on Embedded Multicore/Many-core Systems-on-Chip, MCSoC 2025, Singapore, December 15-18, 2025 (pp. 804-812). Institute of Electrical and Electronics Engineers (IEEE)
Open this publication in new window or tab >>Evaluating Four FPGA-Accelerated Space Use Cases Based on Neural Network Algorithms for On-Board Inference
Show others...
2025 (English)In: Proc. - IEEE Int. Symp. Embed. Multicore/Many-core Syst.-on-Chip, MCSoC, Institute of Electrical and Electronics Engineers (IEEE) , 2025, p. 804-812Conference paper, Published paper (Refereed)
Abstract [en]

Space missions increasingly deploy high-fidelity sensors that produce data volumes exceeding onboard buffering and downlink capacity. This work evaluates FPGA acceleration of neural networks (NNs) across four space use cases on the AMD ZCU104 board. We use Vitis AI (AMD DPU) and Vitis HLS to implement inference, quantify throughput and energy, and expose toolchain and architectural constraints relevant to deployment. Vitis AI achieves up to 34.16 × higher inference rate than the embedded ARM CPU baseline, while custom HLS designs reach up to 5.4 × speedup and add support for operators (e.g., sigmoids, 3D layers) absent in the DPU. For these implementations, measured MPSoC inference power spans 1.56.75 W, reducing energy per inference versus CPU execution in all use cases. These results show that NN FPGA acceleration can enable onboard filtering, compression, and event detection, easing downlink pressure in future missions. 

Place, publisher, year, edition, pages
Institute of Electrical and Electronics Engineers (IEEE), 2025
Keywords
FPGA, HLS, Neural Network, Space Mission, Vitis AI, Embedded systems, Inference engines, Neural networks, Signal processing, Space flight, Case based, Data volume, Energy, High-fidelity, Neural networks algorithms, Neural-networks, Space missions, Space use, Field programmable gate arrays (FPGA)
National Category
Computer Sciences
Identifiers
urn:nbn:se:kth:diva-384585 (URN)10.1109/MCSoC67473.2025.00126 (DOI)2-s2.0-105032402760 (Scopus ID)
Conference
18th IEEE International Symposium on Embedded Multicore/Many-core Systems-on-Chip, MCSoC 2025, Singapore, December 15-18, 2025
Note

Part of ISBN 9798331565718

QC 20260709

Available from: 2026-07-01 Created: 2026-07-01 Last updated: 2026-07-16Bibliographically approved
Organisations
Identifiers
ORCID iD: ORCID iD iconorcid.org/0000-0002-9150-3847

Search in DiVA

Show all publications