Posts by Alessandro Fanfarillo
Performance Profiling on AMD GPUs - Part 6: Advanced Thread Trace (ATT) - The Microscope for Your Application
- 22 September 2026
This post is the sixth part of the Performance Profiling on AMD GPUs series, which works through the
ROCm profiling tools one layer at a time. Part 1 laid the
foundations and introduced the profiling tools available for AMD GPUs,
Parts 2 and 3 walked through
basic and
advanced
profiling workflows with rocprofv3, rocprof-sys, and rocprof-compute, Part 4 applied those
workflows to a
Fortran OpenMP offload application,
and Part 5 used the resulting profiles to drive an
AI-assisted kernel optimization loop.
Those parts tell us which kernel to look at and how it uses the hardware. This part goes one level
below them: Advanced Thread Trace (ATT) records what the wavefronts of a single kernel do
instruction by instruction, so we can explain why that kernel performs the way it does.
Performance Profiling on AMD GPUs - Part 3: Advanced Usage
- 23 October 2025
Error parsing meta tag attribute “keywords”: No content.
Performance Profiling on AMD GPUs – Part 2: Basic Usage
- 13 August 2025
Error parsing meta tag attribute “keywords”: No content.
Performance Profiling on AMD GPUs – Part 1: Foundations
- 26 June 2025
Error parsing meta tag attribute “keywords”: No content.
C++17 parallel algorithms and HIPSTDPAR
- 18 April 2024
The C++17 standard added the concept of parallel algorithms to the
pre-existing C++ Standard Library. The parallel version of algorithms like
std::transform maintain the same signature as the regular serial version,
except for the addition of an extra parameter specifying the
execution policy to use. This flexibility allows users that are already
using the C++ Standard Library algorithms to take advantage of multi-core
architectures by just introducing minimal changes to their code.
Register pressure in AMD CDNA™2 GPUs
- 17 May 2023
Register pressure in GPU kernels has a tremendous impact on the overall performance of your HPC application. Understanding and controlling register usage allows developers to carefully design codes capable of maximizing hardware resources. The following blog post is focused on a practical demo showing how to apply the recommendations explained in this OLCF training talk presented on August 23rd 2022. Here is the training archive where you can also find the slides. We focus solely on the AMD CDNA™2 architecture (MI200 series GPUs) using ROCm 5.4.