kth.sePublikationer KTH
Ändra sökning
Länk till posten
Permanent länk

Direktlänk
Publikationer (10 of 47) Visa alla publikationer
Eleftherakis, P.-E., Anagnostopoulos, G., Kapetanakis, A., Umair, M., Vet, J.-Y., Iliakis, K., . . . Xydis, S. (2026). Multi-Partner Project: Multi-GPU Performance Portability Analysis for CFD Simulations at Scale. In: 2026 Design, Automation and Test in Europe Conference, DATE 2026 - Proceedings: . Paper presented at 2026 Design, Automation and Test in Europe Conference, DATE 2026, Verona, Italy, April 20-22, 2026. Institute of Electrical and Electronics Engineers (IEEE)
Öppna denna publikation i ny flik eller fönster >>Multi-Partner Project: Multi-GPU Performance Portability Analysis for CFD Simulations at Scale
Visa övriga...
2026 (Engelska)Ingår i: 2026 Design, Automation and Test in Europe Conference, DATE 2026 - Proceedings, Institute of Electrical and Electronics Engineers (IEEE) , 2026Konferensbidrag, Publicerat paper (Refereegranskat)
Abstract [en]

As heterogeneous supercomputing architectures leveraging GPUs become increasingly central to high-performance computing (HPC), it is crucial for computational fluid dynamics (CFD) simulations, a de-facto HPC workload, to efficiently utilize such hardware. One of the key challenges of HPC codes is performance portability, i.e. the ability to maintain near-optimal performance across different accelerators. In the context of the REFMAP project, which targets scalable, GPU-enabled multi-fidelity CFD for urban airflow prediction, this paper analyzes the performance portability of SOD2D, a state-of-the-art Spectral Elements simulation framework across AMD and NVIDIA GPU architectures. We first discuss the physical and numerical models underlying SOD2D, highlighting its computational hotspots. Then, we examine its performance and scalability in a multi-level manner, i.e. defining and characterizing an extensive full-stack design space spanning across application, software and hardware infrastructure related parameters. Single-GPU performance characterization across server-grade NVIDIA and AMD GPU architectures and vendor-specific compiler stacks, show the potential as well as the diverse effect of memory access optimizations, i.e. 0.69× - 3.91× deviations in acceleration speedup. Performance variability of SOD2D at scale is further examined on the LUMI multi-GPU cluster, where profiling reveals similar throughput variations, highlighting the limits of performance projections and the need for multi-level, informed tuning.

Ort, förlag, år, upplaga, sidor
Institute of Electrical and Electronics Engineers (IEEE), 2026
Nyckelord
CFD, Performance portability, Spectral Finite Element Method (FEM), design space exploration, high-fidelity simulation, multi-GPU acceleration, scalability analysis
Nationell ämneskategori
Datavetenskap (datalogi) Datorsystem
Identifikatorer
urn:nbn:se:kth:diva-384140 (URN)10.23919/DATE69613.2026.11539345 (DOI)2-s2.0-105041993384 (Scopus ID)
Konferens
2026 Design, Automation and Test in Europe Conference, DATE 2026, Verona, Italy, April 20-22, 2026
Anmärkning

Part of ISBN 9783982674117

QC 20260625

Tillgänglig från: 2026-06-25 Skapad: 2026-06-25 Senast uppdaterad: 2026-06-25Bibliografiskt granskad
Eleftherakis, P.-E., Anagnostopoulos, G., Kapetanakis, A., Umair, M., Vet, J.-Y., Iliakis, K., . . . Xydis, S. (2025). POSTER: Performance Portability in GPU-Accelerated Spectral Finite Element Fluid Simulations: A Cross-layer Exploration Approach. In: Proceedings Of The22Nd Acm International Conference On Computing Frontiers 2025,  CF 2025: . Paper presented at 22nd International Conference on Computing Frontiers-CF, MAY 28-30, 2025, Cagliari, ITALY (pp. 228-229). Association for Computing Machinery (ACM)
Öppna denna publikation i ny flik eller fönster >>POSTER: Performance Portability in GPU-Accelerated Spectral Finite Element Fluid Simulations: A Cross-layer Exploration Approach
Visa övriga...
2025 (Engelska)Ingår i: Proceedings Of The22Nd Acm International Conference On Computing Frontiers 2025,  CF 2025, Association for Computing Machinery (ACM) , 2025, s. 228-229Konferensbidrag, Publicerat paper (Refereegranskat)
Abstract [en]

As heterogeneous supercomputing architectures leveraging GPUs become increasingly central to high-performance computing (HPC), it is crucial for computational fluid dynamics (CFD) simulations to maintain performance portability. In this paper, we examine the performance and scalability of CFD framework SOD2D in a crosslayer manner, i.e. across application, software and hardware infrastructure related parameters. Single-GPU performance characterization across server-grade NVIDIA and AMD GPU architectures and vendor-specific compiler stacks, show the potential as well as the diverse effect of memory access optimizations, i.e. 0.69x - 3.96x deviations in acceleration speedup. Performance variability of SOD2D at scale is then further examined on the LUMI multi-GPU cluster, showcasing analogous diverse effects on throughput, demonstrating the ineffectiveness of adopting performance projections, thus underscoring the importance and necessity of cross-layer informed performance analysis and tuning for multi-GPU configurations.

Ort, förlag, år, upplaga, sidor
Association for Computing Machinery (ACM), 2025
Nyckelord
Performance portability, CFD, Spectral Finite Element Method (FEM), high-fidelity simulation, multi-GPU acceleration, design space exploration, scalability analysis
Nationell ämneskategori
Datavetenskap (datalogi)
Identifikatorer
urn:nbn:se:kth:diva-374072 (URN)10.1145/3719276.3729458 (DOI)001539185100038 ()2-s2.0-105014909534 (Scopus ID)979-8-4007-1528-0 (ISBN)
Konferens
22nd International Conference on Computing Frontiers-CF, MAY 28-30, 2025, Cagliari, ITALY
Anmärkning

QC 20251212

Tillgänglig från: 2025-12-12 Skapad: 2025-12-12 Senast uppdaterad: 2025-12-12Bibliografiskt granskad
D’Orto, M., Sjöblom, S., Chien, L. S., Axner, L. & Gong, J. (2021). Comparing Different Approaches for Solving Large Scale Power-flow Problems with the Newton-Raphson Method. IEEE Access, 9, 56604-56615
Öppna denna publikation i ny flik eller fönster >>Comparing Different Approaches for Solving Large Scale Power-flow Problems with the Newton-Raphson Method
Visa övriga...
2021 (Engelska)Ingår i: IEEE Access, E-ISSN 2169-3536, Vol. 9, s. 56604-56615Artikel i tidskrift (Refereegranskat) Published
Abstract [en]

This paper focuses on using the Newton-Raphson method to solve the power-flow problems. Since the most computationally demanding part of the Newton-Raphson method is to solve the linear equations at each iteration, this study investigates different approaches to solve the linear equations on both central processing unit (CPU) and graphical processing unit (GPU). Six different approaches have been developed and evaluated in this paper: two approaches of these run entirely on CPU while other two of these run entirely on GPU, and the remaining two are hybrid approaches that run on both CPU and GPU. All six direct linear solvers use either LU or QR factorization to solve the linear equations. Two different hardware platforms have been used to conduct the experiments. The performance results show that the CPU version with LU factorization gives better performance compared to the GPU version using standard library called cuSOLVER even for the larger power-flow problems. Moreover, it has been proven that the best performance is achieved using a hybrid method where the Jacobian matrix is assembled on GPU, the preprocessing with a sparse high performance linear solver called KLU is performed on the CPU in the first iteration, and the linear equation is factorized on the GPU and solved on the CPU. Maximum speed up in this study is obtained on the largest case with 25000 buses. The hybrid version shows a speedup factor of 9.6 with a NVIDIA P100 GPU while 13.1 with a NVIDIA V100 GPU in comparison with baseline CPU version on an Intel Xeon Gold 6132 CPU.

Ort, förlag, år, upplaga, sidor
Institute of Electrical and Electronics Engineers (IEEE), 2021
Nationell ämneskategori
Beräkningsmatematik Annan elektroteknik och elektronik
Identifikatorer
urn:nbn:se:kth:diva-292698 (URN)10.1109/ACCESS.2021.3072338 (DOI)000641940300001 ()2-s2.0-85104180269 (Scopus ID)
Anmärkning

QC 20210427

Tillgänglig från: 2021-04-17 Skapad: 2021-04-17 Senast uppdaterad: 2024-03-15Bibliografiskt granskad
Zhang, M., Gong, J., Axner, L. & Barth, M. (2020). Automation of High-Fidelity CFD Analysis for Aircraft Design and Optimization Aided by HPC. In: Proceeding of 28th Euromicro International Conference on Parallel, Distributed and Network-Based Processing (PDP): . Paper presented at 28th Euromicro International Conference on Parallel, Distributed and Network-Based Processing, PDP 2020, Västerås, Sweden, March 11-13, 2020 (pp. 395-399). Institute of Electrical and Electronics Engineers (IEEE)
Öppna denna publikation i ny flik eller fönster >>Automation of High-Fidelity CFD Analysis for Aircraft Design and Optimization Aided by HPC
2020 (Engelska)Ingår i: Proceeding of 28th Euromicro International Conference on Parallel, Distributed and Network-Based Processing (PDP), Institute of Electrical and Electronics Engineers (IEEE) , 2020, s. 395-399Konferensbidrag, Publicerat paper (Refereegranskat)
Abstract [en]

In this paper, an automation process to perform Reynolds-Averaged Navier-Stokes (RANS) computational fluid dynamics (CFD) analysis is developed to carry out aerodynamic design and optimization. The aircraft model/geometry is defined by a Common Parametric Aircraft Configuration Schema (CPACS) file, and the analyses are facilitated using high performance computers (HPC). As the computational capability of the available HPC systems is a limiting factor in the complexity of analyses that can be performed, a detailed performance analysis of the open source CFD code SU2 is undertaken and the profiling and performance analyses for large simulations are carried out.

Ort, förlag, år, upplaga, sidor
Institute of Electrical and Electronics Engineers (IEEE), 2020
Nationell ämneskategori
Datorsystem
Identifikatorer
urn:nbn:se:kth:diva-276248 (URN)10.1109/PDP50117.2020.00067 (DOI)000582555800060 ()2-s2.0-85085483221 (Scopus ID)
Konferens
28th Euromicro International Conference on Parallel, Distributed and Network-Based Processing, PDP 2020, Västerås, Sweden, March 11-13, 2020
Anmärkning

QC 20200610

Tillgänglig från: 2020-06-10 Skapad: 2020-06-10 Senast uppdaterad: 2023-03-30Bibliografiskt granskad
Marco, K., Gong, J., Axner, L., Laure, E. & Jan, N. (2020). GPU-acceleration of A High Order Finite Difference Code Using Curvilinear Coordinates. In: Proceedings of the 2020 International Conference on Computing, Networks and Internet of Things: . Paper presented at the 2020 International Conference on Computing, Networks and Internet of Things (pp. 41-47). Association for Computing Machinery (ACM)
Öppna denna publikation i ny flik eller fönster >>GPU-acceleration of A High Order Finite Difference Code Using Curvilinear Coordinates
Visa övriga...
2020 (Engelska)Ingår i: Proceedings of the 2020 International Conference on Computing, Networks and Internet of Things, Association for Computing Machinery (ACM) , 2020, s. 41-47Konferensbidrag, Publicerat paper (Refereegranskat)
Abstract [en]

GPU-accelerated computing is becoming a popular technology due to the emergence of techniques such as OpenACC, which makes it easy to port codes in their original form to GPU systems using compiler directives, and thereby speeding up computation times relatively simply. In this study we have developed an OpenACC implementation of the high order finite difference CFD solver ESSENSE for simulating compressible flows. The solver is based on summation-by-part form difference operators, and the boundary and interface conditions are weakly implemented using simultaneous approximation terms. This case study focuses on porting code to GPUs for the most time-consuming parts namely sparse matrix vector multiplications and the evaluations of fluxes. The resulting OpenACC implementation is used to simulate the Taylor-Green vortex which produces a maximum speed-up of 61.3 on a single V100 GPU by compared to serial CPU version.

Ort, förlag, år, upplaga, sidor
Association for Computing Machinery (ACM), 2020
Nyckelord
Computational fluid dynamics, GPU programming, High order finite difference method, OpenACC
Nationell ämneskategori
Datorsystem
Identifikatorer
urn:nbn:se:kth:diva-273805 (URN)10.1145/3398329.3398336 (DOI)2-s2.0-85086223863 (Scopus ID)
Konferens
the 2020 International Conference on Computing, Networks and Internet of Things
Anmärkning

QC 20200819

Tillgänglig från: 2020-06-26 Skapad: 2020-06-26 Senast uppdaterad: 2023-03-30Bibliografiskt granskad
Zhang, M., Gong, J. & Axner, L. (2020). HPC-Enabled Aerodynamic Optimization Studies Using CFD and Design Suite SU2. In: Proceeding of the Work in Progress Session held in connection with the PDP 2020 Parallel, Distributed, and Network-Based Processing: . Paper presented at PDP.
Öppna denna publikation i ny flik eller fönster >>HPC-Enabled Aerodynamic Optimization Studies Using CFD and Design Suite SU2
2020 (Engelska)Ingår i: Proceeding of the Work in Progress Session held in connection with the PDP 2020 Parallel, Distributed, and Network-Based Processing, 2020Konferensbidrag, Publicerat paper (Refereegranskat)
Nationell ämneskategori
Datorteknik
Identifikatorer
urn:nbn:se:kth:diva-271128 (URN)
Konferens
PDP
Anmärkning

QC 20200529

Tillgänglig från: 2020-03-18 Skapad: 2020-03-18 Senast uppdaterad: 2024-03-15Bibliografiskt granskad
Otero, E., Gong, J., Min, M., Fischer, P., Schlatter, P. & Laure, E. (2019). OpenACC acceleration for the PN-PN-2 algorithm in Nek5000. Journal of Parallel and Distributed Computing, 132, 69-78
Öppna denna publikation i ny flik eller fönster >>OpenACC acceleration for the PN-PN-2 algorithm in Nek5000
Visa övriga...
2019 (Engelska)Ingår i: Journal of Parallel and Distributed Computing, ISSN 0743-7315, E-ISSN 1096-0848, Vol. 132, s. 69-78Artikel i tidskrift (Refereegranskat) Published
Abstract [en]

Due to its high performance and throughput capabilities, GPU-accelerated computing is becoming a popular technology in scientific computing, in particular using programming models such as CUDA and OpenACC. The main advantage with OpenACC is that it enables to simply port codes in their "original" form to GPU systems through compiler directives, thus allowing an incremental approach. An OpenACC implementation is applied to the CFD code Nek5000 for simulation of incompressible flows, based on the spectral-element method. The work follows up previous implementations and focuses now on the P-N-PN-2 method for the spatial discretization of the Navier-Stokes equations. Performance results of the ported code show a speed-up of up to 3.1 on multi-GPU for a polynomial order N > 11.

Ort, förlag, år, upplaga, sidor
Academic Press, 2019
Nyckelord
Nek5000; OpenACC; GPU programming; Spectral element method; High performance computing
Nationell ämneskategori
Data- och informationsvetenskap
Identifikatorer
urn:nbn:se:kth:diva-253811 (URN)10.1016/j.jpdc.2019.05.010 (DOI)000476580400006 ()2-s2.0-85066835225 (Scopus ID)
Forskningsfinansiär
EU, Horisont 2020Swedish e‐Science Research CenterStiftelsen för strategisk forskning (SSF)
Anmärkning

QC 20190625

Tillgänglig från: 2019-06-18 Skapad: 2019-06-18 Senast uppdaterad: 2022-06-26Bibliografiskt granskad
Zhang, M., Gong, J., Axner, L. & Barth, M. (2019). PRACE Project Airinnova: Automation of High-Fidelity CFD Analysis for Aircraft Design and Optimization. PRACE
Öppna denna publikation i ny flik eller fönster >>PRACE Project Airinnova: Automation of High-Fidelity CFD Analysis for Aircraft Design and Optimization
2019 (Engelska)Rapport (Övrigt vetenskapligt)
Abstract [en]

Airinnova is a start-up company with a key competency in the automation of high-fidelity computational fluid dynamics (CFD) analysis. Following on from our previous PRACE SHAPE project, we have continued collaborating with the PDC Center for High Performance Computing at the KTH Royal Institute of Technology (KTH-PDC), to investigate the performance analysis of the open source CFD code SU2 and further develop the automation process for the field of aerodynamic optimization and design.

Ort, förlag, år, upplaga, sidor
PRACE, 2019. s. 9
Serie
PRACE Whitepaper
Nyckelord
CFD, code, automation process
Nationell ämneskategori
Strömningsmekanik
Identifikatorer
urn:nbn:se:kth:diva-385486 (URN)10.5281/ZENODO.2633710 (DOI)
Projekt
PRACE 5IP - PRACE 5th Implementation Phase Project
Anmärkning

This is a White Paper from a small and medium enterprise (SME) who took part in the PRACE SHAPE Project, funded by PRACE 5IP.

QC 20260717

Tillgänglig från: 2026-07-16 Skapad: 2026-07-16 Senast uppdaterad: 2026-07-17Bibliografiskt granskad
Eliasson, P., Gong, J. & Nordström, J. (2018). A stable and conservative coupling of the unsteady compressible navier-stokes equations at interfaces using finite difference and finite volume methods. In: AIAA Aerospace Sciences Meeting, 2018: . Paper presented at AIAA Aerospace Sciences Meeting, 2018, Kissimmee, United States, 8 January 2018 through 12 January 2018. American Institute of Aeronautics and Astronautics Inc, AIAA (210059)
Öppna denna publikation i ny flik eller fönster >>A stable and conservative coupling of the unsteady compressible navier-stokes equations at interfaces using finite difference and finite volume methods
2018 (Engelska)Ingår i: AIAA Aerospace Sciences Meeting, 2018, American Institute of Aeronautics and Astronautics Inc, AIAA , 2018, nr 210059Konferensbidrag, Publicerat paper (Refereegranskat)
Abstract [en]

Stable and conservative interface boundary conditions are developed for the unsteady compressible Navier-Stokes equations using finite difference and finite volume methods. The finite difference approach is based on summation-by-part operators and can be made higher order accurate with boundary conditions imposed weakly. The finite volume approach is an edge- and dual grid-based approach for unstructured grids, formally second order accurate in space, with weak boundary conditions as well. Stable and conservative weak boundary conditions are derived for interfaces between finite difference methods, for finite volume methods and for the coupling between the two approaches. The three types of interface boundary conditions are demonstrated for two test cases. Firstly, inviscid vortex propagation with a known analytical solution is considered. The results show expected error decays as the grid is refined for various couplings and spatial accuracy of the finite difference scheme. The second test case involves viscous laminar flow over a cylinder with vortex shedding. Calculations with various coupling and spatial accuracies of the finite difference solver show that the couplings work as expected and that the higher order finite difference schemes provide enhanced vortex propagation.

Ort, förlag, år, upplaga, sidor
American Institute of Aeronautics and Astronautics Inc, AIAA, 2018
Nationell ämneskategori
Beräkningsmatematik
Identifikatorer
urn:nbn:se:kth:diva-225496 (URN)10.2514/6.2018-0597 (DOI)2-s2.0-85141554093 (Scopus ID)9781624105241 (ISBN)
Konferens
AIAA Aerospace Sciences Meeting, 2018, Kissimmee, United States, 8 January 2018 through 12 January 2018
Forskningsfinansiär
Swedish e‐Science Research CenterStiftelsen för internationalisering av högre utbildning och forskning (STINT)VINNOVA
Anmärkning

QC 20180406

Tillgänglig från: 2018-04-06 Skapad: 2018-04-06 Senast uppdaterad: 2023-06-08Bibliografiskt granskad
Larsson, T., Hammar, J., Gong, J., Barth, M. & Axner, L. (2018). ENHANCING COMPUTATIONAL AERO-ACOUSTIC PROCESSES FOR GROUNDVEHICLES RESOLVING OPEN SOURCE CFD. In: The 13th OpenFOAM Workshop: . Paper presented at The 13th OpenFOAM Workshop (pp. 1-4).
Öppna denna publikation i ny flik eller fönster >>ENHANCING COMPUTATIONAL AERO-ACOUSTIC PROCESSES FOR GROUNDVEHICLES RESOLVING OPEN SOURCE CFD
Visa övriga...
2018 (Engelska)Ingår i: The 13th OpenFOAM Workshop, 2018, s. 1-4Konferensbidrag, Muntlig presentation med publicerat abstract (Refereegranskat)
Nationell ämneskategori
Strömningsmekanik
Identifikatorer
urn:nbn:se:kth:diva-232361 (URN)
Konferens
The 13th OpenFOAM Workshop
Anmärkning

QC 20180821

Tillgänglig från: 2018-07-20 Skapad: 2018-07-20 Senast uppdaterad: 2025-02-09Bibliografiskt granskad
Organisationer
Identifikatorer
ORCID-id: ORCID iD iconorcid.org/0000-0002-3859-9480

Sök vidare i DiVA

Visa alla publikationer