kth.sePublications KTH
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Stencil Computations on AMD and Nvidia Graphics Processors: Performance and Tuning Strategies
Aalto Univ, Dept Comp Sci, Espoo, Finland.
Univ Helsinki, Dept Comp Sci, Helsinki, Finland.
CSC IT Ctr Sci Ltd, Espoo, Finland.
Aalto Univ, Dept Comp Sci, Espoo, Finland; Max Planck Inst Solar Syst Res, Gottingen, Germany; KTH Royal Inst Technol, Nordita, 10691 Stockholm, Sweden; Stockholm Univ, Stockholm, Sweden.
2025 (English)In: Concurrency and Computation, ISSN 1532-0626, E-ISSN 1532-0634, Vol. 37, no 12-14, article id e70129Article in journal (Refereed) Published
Abstract [en]

Over the last ten years, graphics processors have become the de facto accelerator for data-parallel tasks in various branches of high-performance computing, including machine learning and computational sciences. However, with the recent introduction of AMD-manufactured graphics processors to the world's fastest supercomputers, tuning strategies established for previous hardware generations must be re-evaluated. In this study, we evaluate the performance and energy efficiency of stencil computations on modern datacenter graphics processors and propose a tuning strategy for fusing cache-heavy stencil kernels. The studied cases comprise both synthetic and practical applications, which involve the evaluation of linear and nonlinear stencil functions in one to three dimensions. Our experiments reveal that AMD and Nvidia graphics processors exhibit key differences in both hardware and software, necessitating platform-specific tuning to reach their full computational potential.

Place, publisher, year, edition, pages
Wiley , 2025. Vol. 37, no 12-14, article id e70129
Keywords [en]
discrete convolution, energy efficiency, graphics processing units, high-performance computing, partial differential equations, performance optimization, stencil computations
National Category
Computer Sciences
Identifiers
URN: urn:nbn:se:kth:diva-367916DOI: 10.1002/cpe.70129ISI: 001497532600026Scopus ID: 2-s2.0-105006774862OAI: oai:DiVA.org:kth-367916DiVA, id: diva2:1987524
Note

QC 20250806

Available from: 2025-08-06 Created: 2025-08-06 Last updated: 2025-08-06Bibliographically approved

Open Access in DiVA

No full text in DiVA

Other links

Publisher's full textScopus

Search in DiVA

By author/editor
Korpi-Lagg, Maarit J.
In the same journal
Concurrency and Computation
Computer Sciences

Search outside of DiVA

GoogleGoogle Scholar

doi
urn-nbn

Altmetric score

doi
urn-nbn
Total: 24 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf