kth.sePublications KTH
Change search
Link to record
Permanent link

Direct link
Publications (5 of 5) Show all publications
Pennati, L., Ekelund, J., Hu, A., Peng, I. & Markidis, S. (2026). Disturbance storm time index prediction with interpretable machine learning. Journal of Computational Science, 95, Article ID 102821.
Open this publication in new window or tab >>Disturbance storm time index prediction with interpretable machine learning
Show others...
2026 (English)In: Journal of Computational Science, ISSN 1877-7503, E-ISSN 1877-7511, Vol. 95, article id 102821Article in journal (Refereed) Published
Abstract [en]

The Disturbance Storm Time (Dst) index quantifies geomagnetic storm intensity by measuring global magnetic field variations. In this study, we apply interpretable machine-learning (ML) techniques to derive data-driven models describing the temporal evolution of the Dst index. We use historical data from the NASA OMNIWeb database, including solar wind density, bulk velocity, convective electric field, dynamic pressure, and magnetic pressure. We employ KAN networks and the symbolic regression framework PyOperon, based on an evolutionary algorithm, to identify closed-form expressions linking (Formula presented) to key solar wind parameters. The equations obtained via symbolic regression form a hierarchy of complexity levels and capture nonlinear dependencies and threshold effects in Dst evolution. In addition, we use a conventional MLP network as a reference black-box model. We benchmark all ML models against observed Dst data and compare their performance with empirical formulations such as the Burton-McPherron–Russell and O’Brien-McPherron models. The performance evaluation on historical storm events includes the 2003 Halloween storm, the 2015 St. Patrick’s Day storm, a moderate storm in 2017, and the extreme storm of May 2024. The data-driven models, particularly the MLP, demonstrate superior accuracy in most cases. While the symbolic regression expressions provide insight into the underlying physics, the results highlight an intrinsic trade-off between model interpretability and predictive accuracy. This is an extended version of a previous work presented in Markidis et al. (2025) [1].

Place, publisher, year, edition, pages
Elsevier BV, 2026
Keywords
Dst index prediction, Geomagnetic storms, Interpretable machine learning, Symbolic regression
National Category
Astronomy, Astrophysics and Cosmology Other Computer and Information Science Bioinformatics (Computational Biology)
Identifiers
urn:nbn:se:kth:diva-379284 (URN)10.1016/j.jocs.2026.102821 (DOI)001709589400001 ()2-s2.0-105034374796 (Scopus ID)
Note

QC 20260417

Available from: 2026-04-17 Created: 2026-04-17 Last updated: 2026-04-17Bibliographically approved
Hübner, P., Hu, A., Peng, I. & Markidis, S. (2025). Apple vs. Oranges: Evaluating the Apple Silicon M-Series SoCs for HPC Performance and Efficiency. In: Proceedings - 2025 IEEE International Parallel and Distributed Processing Symposium Workshops, IPDPSW 2025: . Paper presented at 2025 IEEE International Parallel and Distributed Processing Symposium Workshops, IPDPSW 2025, Milan, Italy, June 3-7, 2025 (pp. 45-54). Institute of Electrical and Electronics Engineers (IEEE)
Open this publication in new window or tab >>Apple vs. Oranges: Evaluating the Apple Silicon M-Series SoCs for HPC Performance and Efficiency
2025 (English)In: Proceedings - 2025 IEEE International Parallel and Distributed Processing Symposium Workshops, IPDPSW 2025, Institute of Electrical and Electronics Engineers (IEEE) , 2025, p. 45-54Conference paper, Published paper (Refereed)
Abstract [en]

This paper investigates the architectural features and performance potential of the Apple Silicon M-Series SoCs (M1, M2, M3, and M4) for HPC. We provide a detailed review of the CPU and GPU designs, the unified memory architecture, and coprocessors such as Advanced Matrix Extensions (AMX). We design and develop benchmarks in the Metal Shading Language and Objective-C++ to assess FP32 computational and memory performance. We also measure power consumption and efficiency using Apple's powermetrics tool. Our results show that the M-Series chips offer up to 100 GB/s memory bandwidth, and significant generational improvements in computational performance, with up to 2.9 FP32 TFLOPS on the M4. Power consumption varies from a few Watts to 10-20 Watts, with more than 200 GFLOPS per Watt efficiency of GPU and accelerator reached by all four chips. Despite limitations in FP64 support on the GPU, the M-Series chips demonstrate strong potential for energy-efficient HPC applications. While existing HPC solutions such as the Nvidia Grace-Hopper superchip outperform Apple Silicon in both memory bandwidth and computational performance, we see that the M-Series provides a competitive power-efficient alternative to traditional HPC architectures and represents a distinct category altogether - forming an apples-to-oranges comparison.

Place, publisher, year, edition, pages
Institute of Electrical and Electronics Engineers (IEEE), 2025
Keywords
Apple Silicon M-Series GPU Performance, ARM-based SoC, M1, M2, M3, M4 Architecture
National Category
Computer Systems
Identifiers
urn:nbn:se:kth:diva-370765 (URN)10.1109/IPDPSW66978.2025.00013 (DOI)001566005900005 ()2-s2.0-105015528421 (Scopus ID)
Conference
2025 IEEE International Parallel and Distributed Processing Symposium Workshops, IPDPSW 2025, Milan, Italy, June 3-7, 2025
Note

Part of ISBN 9798331526436

QC 20251001

Available from: 2025-10-01 Created: 2025-10-01 Last updated: 2026-05-29Bibliographically approved
Lumsden, I., Devarajan, H., Yildirim, I., Markidis, S., Hu, A., Peng, I., . . . Taufer, M. (2025). Optimizing I/O for an Exascale Implicit Kinetic Plasma Simulation using the Rabbit Storage System. In: 2025 IEEE International Conference On Cluster Computing Workshops, Cluster Workshops: . Paper presented at 2025 International Conference on Cluster Computing-CLUSTER-Annual, SEP 02-05, 2025, University of Edinburgh, Edinburgh, ENGLAND (pp. 61-66). Institute of Electrical and Electronics Engineers (IEEE)
Open this publication in new window or tab >>Optimizing I/O for an Exascale Implicit Kinetic Plasma Simulation using the Rabbit Storage System
Show others...
2025 (English)In: 2025 IEEE International Conference On Cluster Computing Workshops, Cluster Workshops, Institute of Electrical and Electronics Engineers (IEEE) , 2025, p. 61-66Conference paper, Published paper (Refereed)
Abstract [en]

Exascale computing revolutionizes scientific research by facilitating large-scale simulations and data analysis. As applications generate massive, heterogeneous data streams, the I/O subsystem increasingly dominates runtime, limiting scalability and time-to-solution. Emerging storage accelerators such as the Rabbit system offer a promising path forward. Rabbit provides dynamically configurable hybrid storage, combining high-bandwidth, node-local NVMe (e.g., XFS) with shared, distributed file systems (e.g., Lustre), to match diverse I/O patterns more efficiently. In this work, we evaluate Rabbit's capability to mitigate I/O bottlenecks using iPIC3D, a representative implicit Particle-in-Cell (PIC) code for exascale scientific workloads. We perform a detailed characterization of the key phases of I/O (restart, field, and moment) and identify the dominant access patterns in these phases. We then benchmark these patterns using IOR in Rabbit hybrid storage configurations, identifying the ideal performance achievable for each access pattern under node-local and distributed storage modes. We use phase-aware mapping of these patterns to Rabbit storage systems, achieving a 4.85 times improvement in I/O throughput, reducing the I/O share of runtime from 38% to 11% and delivering an end-to-end speedup of 1.45 times. These results outline the importance of hardware-software co-design and highlight Rabbit as a scalable, data-intensive storage solution.

Place, publisher, year, edition, pages
Institute of Electrical and Electronics Engineers (IEEE), 2025
Keywords
High performance computing, I/O optimization, Rabbit nodes, I/O accelerators, iPIC3D, IOR
National Category
Computer Sciences
Identifiers
urn:nbn:se:kth:diva-381886 (URN)10.1109/CLUSTERWorkshops65972.2025.11164212 (DOI)001704667500019 ()2-s2.0-105018082545 (Scopus ID)
Conference
2025 International Conference on Cluster Computing-CLUSTER-Annual, SEP 02-05, 2025, University of Edinburgh, Edinburgh, ENGLAND
Note

Part of ISBN 9798331512569

QC 20260525

Available from: 2026-05-25 Created: 2026-05-25 Last updated: 2026-07-14Bibliographically approved
Lumsden, I., Markidis, S., Hu, A., Peng, I., Pennati, L., Yokelson, D., . . . Taufer, M. (2025). Performance Optimization of an Exascale Implicit Kinetic Plasma Simulation on El Capitan. In: Proceedings IEEE International Conference on eScience, eScience 2025: . Paper presented at IEEE International Conference on eScience, eScience 2025, Chicago, IL, USA, September 15-18, 2025 (pp. 319-320). Institute of Electrical and Electronics Engineers (IEEE)
Open this publication in new window or tab >>Performance Optimization of an Exascale Implicit Kinetic Plasma Simulation on El Capitan
Show others...
2025 (English)In: Proceedings IEEE International Conference on eScience, eScience 2025, Institute of Electrical and Electronics Engineers (IEEE) , 2025, p. 319-320Conference paper, Published paper (Refereed)
Abstract [en]

We present performance scaling and optimization of iPIC3D - an exascale-class, GPU-enabled implicit particle-in-cell code for planetary-scale magnetosphere modeling and plasma simulation - on El Capitan. Our strong and weak scaling studies demonstrate near-linear scaling up to 8,000 nodes (32,000 APUs) with a parallel efficiency of nearly 100%. We optimize iPIC3D to leverage AMD’s MI300A APUs, the Merced Lustre filesystem, and Rabbit nodes for high-bandwidth I/O. Optimizations reduce memory usage by 97% and runtime by 74%, enabling simulations that are 1.8 times larger than before. Rabbit further improve checkpointing bandwidth by 2 times, ensuring scalable fault- tolerant simulations on exascale architectures.

Place, publisher, year, edition, pages
Institute of Electrical and Electronics Engineers (IEEE), 2025
National Category
Computational Mathematics
Identifiers
urn:nbn:se:kth:diva-382348 (URN)10.1109/eScience65000.2025.00050 (DOI)001710422500011 ()2-s2.0-105019536147 (Scopus ID)
Conference
IEEE International Conference on eScience, eScience 2025, Chicago, IL, USA, September 15-18, 2025
Note

Part of ISBN 9798331591465, 9798331591458

QC 20260526

Available from: 2026-05-26 Created: 2026-05-26 Last updated: 2026-07-14Bibliographically approved
Hu, A., Pennati, L., Peng, I. & Markidis, S. (2025). Physics-Aware Compression of Plasma Distribution Functions with GPU-Accelerated Gaussian Mixture Models. In: Computational Science - ICCS 2025 - 25th International Conference, 2025, Proceedings: . Paper presented at 25th International Conference on Computational Science, ICCS 2025, Singapore, Singapore, Jul 07 2025 - Jul 09 2025 (pp. 33-47). Springer Nature, 15905 LNCS
Open this publication in new window or tab >>Physics-Aware Compression of Plasma Distribution Functions with GPU-Accelerated Gaussian Mixture Models
2025 (English)In: Computational Science - ICCS 2025 - 25th International Conference, 2025, Proceedings, Springer Nature , 2025, Vol. 15905 LNCS, p. 33-47Conference paper, Published paper (Refereed)
Abstract [en]

Data compression is a critical technology for large-scale plasma simulations. Storing complete particle information requires Terabyte-scale data storage, and analysis requires ad-hoc scalable post-processing tools. We propose a physics-aware in-situ compression method using Gaussian Mixture Models (GMMs) to approximate electron and ion velocity distribution functions with a number of Gaussian components. This GMM-based method allows us to capture plasma features such as mean velocity and temperature, and it enables us to identify heating processes and generate beams. We first construct a histogram to reduce computational overhead and apply GPU-accelerated, in-situ GMM fitting within iPIC3D, a large-scale implicit Particle-in-Cell simulator, ensuring real-time compression. The compressed representation is stored using the ADIOS 2 library, thus optimizing the I/O process. The GPU and histogramming implementation provides a significant speed-up with respect to GMM on particles (both in time and required memory at run-time), enabling real-time compression. Compared to algorithms like SZ, MGARD, and BLOSC2, our GMM-based method has a physics-based approach, retaining the physical interpretation of plasma phenomena such as beam formation, acceleration, and heating mechanisms. Our GMM algorithm achieves a compression ratio of up to 104, requiring a processing time comparable to, or even lower than, standard compression engines.

Place, publisher, year, edition, pages
Springer Nature, 2025
Keywords
Compression Particle-in-Cell, Distribution Functions, Gaussian-Mixture-Model Compression
National Category
Computational Mathematics
Identifiers
urn:nbn:se:kth:diva-385679 (URN)10.1007/978-3-031-97632-2_3 (DOI)2-s2.0-105010828825 (Scopus ID)
Conference
25th International Conference on Computational Science, ICCS 2025, Singapore, Singapore, Jul 07 2025 - Jul 09 2025
Note

Part of ISBN 9783031976315

QC 20260721

Available from: 2026-07-21 Created: 2026-07-21 Last updated: 2026-07-21Bibliographically approved
Organisations
Identifiers
ORCID iD: ORCID iD iconorcid.org/0009-0009-8783-8335

Search in DiVA

Show all publications