Memory Prefetching Evaluation of Scientific Applications on a Modern HPC Arm-Based ProcessorShow others and affiliations
2025 (English)In: IEEE Access, E-ISSN 2169-3536, Vol. 13, p. 85898-85926
Article in journal (Refereed) Published
Abstract [en]
Memory prefetching is a well-known technique for mitigating the negative impact of memory access latencies on memory bandwidth. This problem has become more pressing as improvements in memory bandwidth have not kept pace with increases in computational power. While much existing work has been devoted to finding appropriate prefetching techniques for specific workloads, few provide insight into the behavior of scientific applications to better understand the impact of prefetchers. This paper investigates the impact of hardware prefetchers on the latest Arm-based high-end processor architectures. In this work, we investigate memory access patterns by analyzing locality properties and visualizing delta and repetitive address patterns. A deeper understanding of memory access patterns allows the use of the appropriate prefetcher and reaching a better correlation between access pattern properties and prefetcher performance. This can guide future co-design efforts. We evaluated traditional and innovative prefetchers using a gem5-based model of Arm Neoverse V1 cores. The model features a 16-core architecture, using Amazon's Graviton 3 processor as a hardware reference, but substituting DDR5 by high bandwidth memory (HBM2). We performed a detailed prefetching evaluation focusing on stencil, sparse matrix-vector multiplication, and Breadth-First Search kernels. These kernels represent a broad range of the applications running on today's High-Performance Computing (HPC) systems, which are sensitive to memory performance.
Place, publisher, year, edition, pages
Institute of Electrical and Electronics Engineers (IEEE) , 2025. Vol. 13, p. 85898-85926
Keywords [en]
Prefetching, Bandwidth, Kernel, Hardware, Correlation, Sparse matrices, Single instruction multiple data, Memory management, Layout, High performance computing, Memory prefetcher, computer simulation
National Category
Computer Systems
Identifiers
URN: urn:nbn:se:kth:diva-367921DOI: 10.1109/ACCESS.2025.3569533ISI: 001492121500023Scopus ID: 2-s2.0-105005365333OAI: oai:DiVA.org:kth-367921DiVA, id: diva2:1987517
Note
QC 20250806
2025-08-062025-08-062025-08-06Bibliographically approved