kth.sePublications KTH
Change search
Link to record
Permanent link

Direct link
Vincent, Jonathan
Publications (5 of 5) Show all publications
Eleftherakis, P.-E., Anagnostopoulos, G., Kapetanakis, A., Umair, M., Vet, J.-Y., Iliakis, K., . . . Xydis, S. (2026). Multi-Partner Project: Multi-GPU Performance Portability Analysis for CFD Simulations at Scale. In: 2026 Design, Automation and Test in Europe Conference, DATE 2026 - Proceedings: . Paper presented at 2026 Design, Automation and Test in Europe Conference, DATE 2026, Verona, Italy, April 20-22, 2026. Institute of Electrical and Electronics Engineers (IEEE)
Open this publication in new window or tab >>Multi-Partner Project: Multi-GPU Performance Portability Analysis for CFD Simulations at Scale
Show others...
2026 (English)In: 2026 Design, Automation and Test in Europe Conference, DATE 2026 - Proceedings, Institute of Electrical and Electronics Engineers (IEEE) , 2026Conference paper, Published paper (Refereed)
Abstract [en]

As heterogeneous supercomputing architectures leveraging GPUs become increasingly central to high-performance computing (HPC), it is crucial for computational fluid dynamics (CFD) simulations, a de-facto HPC workload, to efficiently utilize such hardware. One of the key challenges of HPC codes is performance portability, i.e. the ability to maintain near-optimal performance across different accelerators. In the context of the REFMAP project, which targets scalable, GPU-enabled multi-fidelity CFD for urban airflow prediction, this paper analyzes the performance portability of SOD2D, a state-of-the-art Spectral Elements simulation framework across AMD and NVIDIA GPU architectures. We first discuss the physical and numerical models underlying SOD2D, highlighting its computational hotspots. Then, we examine its performance and scalability in a multi-level manner, i.e. defining and characterizing an extensive full-stack design space spanning across application, software and hardware infrastructure related parameters. Single-GPU performance characterization across server-grade NVIDIA and AMD GPU architectures and vendor-specific compiler stacks, show the potential as well as the diverse effect of memory access optimizations, i.e. 0.69× - 3.91× deviations in acceleration speedup. Performance variability of SOD2D at scale is further examined on the LUMI multi-GPU cluster, where profiling reveals similar throughput variations, highlighting the limits of performance projections and the need for multi-level, informed tuning.

Place, publisher, year, edition, pages
Institute of Electrical and Electronics Engineers (IEEE), 2026
Keywords
CFD, Performance portability, Spectral Finite Element Method (FEM), design space exploration, high-fidelity simulation, multi-GPU acceleration, scalability analysis
National Category
Computer Sciences Computer Systems
Identifiers
urn:nbn:se:kth:diva-384140 (URN)10.23919/DATE69613.2026.11539345 (DOI)2-s2.0-105041993384 (Scopus ID)
Conference
2026 Design, Automation and Test in Europe Conference, DATE 2026, Verona, Italy, April 20-22, 2026
Note

Part of ISBN 9783982674117

QC 20260625

Available from: 2026-06-25 Created: 2026-06-25 Last updated: 2026-06-25Bibliographically approved
Eleftherakis, P.-E., Anagnostopoulos, G., Kapetanakis, A., Umair, M., Vet, J.-Y., Iliakis, K., . . . Xydis, S. (2025). POSTER: Performance Portability in GPU-Accelerated Spectral Finite Element Fluid Simulations: A Cross-layer Exploration Approach. In: Proceedings Of The22Nd Acm International Conference On Computing Frontiers 2025,  CF 2025: . Paper presented at 22nd International Conference on Computing Frontiers-CF, MAY 28-30, 2025, Cagliari, ITALY (pp. 228-229). Association for Computing Machinery (ACM)
Open this publication in new window or tab >>POSTER: Performance Portability in GPU-Accelerated Spectral Finite Element Fluid Simulations: A Cross-layer Exploration Approach
Show others...
2025 (English)In: Proceedings Of The22Nd Acm International Conference On Computing Frontiers 2025,  CF 2025, Association for Computing Machinery (ACM) , 2025, p. 228-229Conference paper, Published paper (Refereed)
Abstract [en]

As heterogeneous supercomputing architectures leveraging GPUs become increasingly central to high-performance computing (HPC), it is crucial for computational fluid dynamics (CFD) simulations to maintain performance portability. In this paper, we examine the performance and scalability of CFD framework SOD2D in a crosslayer manner, i.e. across application, software and hardware infrastructure related parameters. Single-GPU performance characterization across server-grade NVIDIA and AMD GPU architectures and vendor-specific compiler stacks, show the potential as well as the diverse effect of memory access optimizations, i.e. 0.69x - 3.96x deviations in acceleration speedup. Performance variability of SOD2D at scale is then further examined on the LUMI multi-GPU cluster, showcasing analogous diverse effects on throughput, demonstrating the ineffectiveness of adopting performance projections, thus underscoring the importance and necessity of cross-layer informed performance analysis and tuning for multi-GPU configurations.

Place, publisher, year, edition, pages
Association for Computing Machinery (ACM), 2025
Keywords
Performance portability, CFD, Spectral Finite Element Method (FEM), high-fidelity simulation, multi-GPU acceleration, design space exploration, scalability analysis
National Category
Computer Sciences
Identifiers
urn:nbn:se:kth:diva-374072 (URN)10.1145/3719276.3729458 (DOI)001539185100038 ()2-s2.0-105014909534 (Scopus ID)979-8-4007-1528-0 (ISBN)
Conference
22nd International Conference on Computing Frontiers-CF, MAY 28-30, 2025, Cagliari, ITALY
Note

QC 20251212

Available from: 2025-12-12 Created: 2025-12-12 Last updated: 2025-12-12Bibliographically approved
Vincent, J., Gong, J., Karp, M., Peplinski, A., Jansson, N., Podobas, A., . . . Schlatter, P. (2022). Strong Scaling of OpenACC enabled Nek5000 on several GPU based HPC systems. In: HPCAsia2022: International Conference on High Performance Computing in Asia-Pacific Region. Paper presented at HPC Asia2022: International Conference on High Performance Computing in Asia-Pacific Region Virtual Event Japan January 12 - 14, 2022 (pp. 94-102). Association for Computing Machinery (ACM)
Open this publication in new window or tab >>Strong Scaling of OpenACC enabled Nek5000 on several GPU based HPC systems
Show others...
2022 (English)In: HPCAsia2022: International Conference on High Performance Computing in Asia-Pacific Region, Association for Computing Machinery (ACM) , 2022, p. 94-102Conference paper, Published paper (Refereed)
Abstract [en]

We present new results on the strong parallel scaling for the OpenACC-accelerated implementation of the high-order spectral element fluid dynamics solver Nek5000. The test case considered consists of a direct numerical simulation of fully-developed turbulent flow in a straight pipe, at two different Reynolds numbers Reτ = 360 and Reτ = 550, based on friction velocity and pipe radius. The strong scaling is tested on several GPU-enabled HPC systems, including the Swiss Piz Daint system, TACC's Longhorn, Jülich's JUWELS Booster, and Berzelius in Sweden. The performance results show that speed-up between 3-5 can be achieved using the GPU accelerated version compared with the CPU version on these different systems. The run-time for 20 timesteps reduces from 43.5 to 13.2 seconds with increasing the number of GPUs from 64 to 512 for Reτ = 550 case on JUWELS Booster system. This illustrates the GPU accelerated version the potential for high throughput. At the same time, the strong scaling limit is significantly larger for GPUs, at about 2000 - 5000 elements per rank; compared to about 50 - 100 for a CPU-rank.

Place, publisher, year, edition, pages
Association for Computing Machinery (ACM), 2022
Series
ACM International Conference Proceeding Series
National Category
Computer Sciences Fluid Mechanics
Identifiers
urn:nbn:se:kth:diva-309189 (URN)10.1145/3492805.3492818 (DOI)2-s2.0-85122621284 (Scopus ID)
Conference
HPC Asia2022: International Conference on High Performance Computing in Asia-Pacific Region Virtual Event Japan January 12 - 14, 2022
Note

QC 20220223

Part of conference proceedings: ISBN 978-145038498-8

Available from: 2022-02-22 Created: 2022-02-22 Last updated: 2025-02-09Bibliographically approved
Laure, E., Ahlin, D., Malinowsky, L., Svensson, G. & Vincent, J. (2015). Lindgren the Swedish tier-1 system. In: Contemporary High Performance Computing: From Petascale Toward Exascale: Volume Two: (pp. 141-162). CRC Press
Open this publication in new window or tab >>Lindgren the Swedish tier-1 system
Show others...
2015 (English)In: Contemporary High Performance Computing: From Petascale Toward Exascale: Volume Two, CRC Press , 2015, p. 141-162Chapter in book (Other academic)
Abstract [en]

The Swedish academic computing landscape is organized under the auspices of SNIC, the Swedish National Infrastructure for Computing. SNIC coordinates investments in computing and storage infrastructure at its six national centers and manages the national process for allocating research time on its computing resources. Since its formation in 2003, SNIC has significantly increased the computational capacity available to Swedish researchers and firmly put Sweden on the international computational science map. When the Partnership for Advanced Computing in Europe (PRACE) started in 2010, SNIC joined this European HPC effort and worked with the Swedish Research 142Council to allocate additional funds for a national high-end system that would also be made available to European researchers via PRACE. These efforts resulted in the installation of a CRAY XE6 supercomputer, named Lindgren, at the PDC Center for High-Performance Computing at the KTH Royal Institute of Technology in Stockholm. 

Place, publisher, year, edition, pages
CRC Press, 2015
National Category
Computer Systems
Identifiers
urn:nbn:se:kth:diva-236901 (URN)2-s2.0-85053978122 (Scopus ID)
Note

Part of ISBN 978-149870063-4, 978-149870062-7

QC 20241212

Available from: 2018-12-12 Created: 2018-12-12 Last updated: 2024-12-12Bibliographically approved
Wöhri, A. B., Katona, G., Johansson, L. C., Fritz, E., Malmerberg, E., Andersson, M., . . . Neutze, R. (2010). Light-Induced Structural Changes in a Photosynthetic Reaction Center Caught by Laue Diffraction. Science, 328(5978), 630-633
Open this publication in new window or tab >>Light-Induced Structural Changes in a Photosynthetic Reaction Center Caught by Laue Diffraction
Show others...
2010 (English)In: Science, ISSN 0036-8075, Vol. 328, no 5978, p. 630-633Article in journal (Refereed) Published
Abstract [en]

Photosynthetic reaction centers convert the energy content of light into a transmembrane potential difference and so provide the major pathway for energy input into the biosphere. We applied time-resolved Laue diffraction to study light-induced conformational changes in the photosynthetic reaction center complex of Blastochloris viridis. The side chain of TyrL162, which lies adjacent to the special pair of bacteriochlorophyll molecules that are photooxidized in the primary light conversion event of photosynthesis, was observed to move 1.3 angstroms closer to the special pair after photoactivation. Free energy calculations suggest that this movement results from the deprotonation of this conserved tyrosine residue and provides a mechanism for stabilizing the primary charge separation reactions of photosynthesis.

National Category
Computer and Information Sciences
Identifiers
urn:nbn:se:kth:diva-75455 (URN)10.1126/science.1186159 (DOI)000277159800046 ()20431017 (PubMedID)2-s2.0-77951854791 (Scopus ID)
Note
QC 20120207Available from: 2012-02-05 Created: 2012-02-05 Last updated: 2024-03-18Bibliographically approved
Organisations

Search in DiVA

Show all publications