Posts by Alessandro Fanfarillo

Performance Profiling on AMD GPUs - Part 6: Advanced Thread Trace (ATT) - The Microscope for Your Application

This post is the sixth part of the Performance Profiling on AMD GPUs series, which works through the ROCm profiling tools one layer at a time. Part 1 laid the foundations and introduced the profiling tools available for AMD GPUs, Parts 2 and 3 walked through basic and advanced profiling workflows with rocprofv3, rocprof-sys, and rocprof-compute, Part 4 applied those workflows to a Fortran OpenMP offload application, and Part 5 used the resulting profiles to drive an AI-assisted kernel optimization loop. Those parts tell us which kernel to look at and how it uses the hardware. This part goes one level below them: Advanced Thread Trace (ATT) records what the wavefronts of a single kernel do instruction by instruction, so we can explain why that kernel performs the way it does.

Read more ...


Performance Profiling on AMD GPUs - Part 3: Advanced Usage

Error parsing meta tag attribute “keywords”: No content.

Read more ...


Performance Profiling on AMD GPUs – Part 2: Basic Usage

Error parsing meta tag attribute “keywords”: No content.

Read more ...


Performance Profiling on AMD GPUs – Part 1: Foundations

Error parsing meta tag attribute “keywords”: No content.

Read more ...


C++17 parallel algorithms and HIPSTDPAR

The C++17 standard added the concept of parallel algorithms to the pre-existing C++ Standard Library. The parallel version of algorithms like std::transform maintain the same signature as the regular serial version, except for the addition of an extra parameter specifying the execution policy to use. This flexibility allows users that are already using the C++ Standard Library algorithms to take advantage of multi-core architectures by just introducing minimal changes to their code.

Read more ...


Register pressure in AMD CDNA™2 GPUs

Register pressure in GPU kernels has a tremendous impact on the overall performance of your HPC application. Understanding and controlling register usage allows developers to carefully design codes capable of maximizing hardware resources. The following blog post is focused on a practical demo showing how to apply the recommendations explained in this OLCF training talk presented on August 23rd 2022. Here is the training archive where you can also find the slides. We focus solely on the AMD CDNA™2 architecture (MI200 series GPUs) using ROCm 5.4.

Read more ...