Posts by Yuchen Lin
Serving GLM-5.2-MXFP4 on AMD Instinct™ MI355X: When Prefill Context Parallelism Pays
- 21 September 2026
Whether prefill context parallelism pays depends on prompt length. It divides the prompt across GPUs during prefill; on four AMD Instinct™ MI355X GPUs (gfx950) serving GLM-5.2-MXFP4 under SGLang at --tp-size 4, enabling it yields 43% to 54% more total throughput on 40,960- and 61,440-token prompts at concurrency 16 and above, and a first token about 1.5× to 2.3× faster at every concurrency. On 1024-token prompts it loses at every concurrency measured, by 8% to 18% of total throughput.
Iteratively Tuning hipBLASLt TensileLite Kernels: A Smaller Search, a Faster Kernel
- 09 September 2026
A TensileLite run can only return a kernel its configuration’s search space can express. That sounds like a technicality, but it is the one fact that decides whether a tuning campaign helps or hurts. The configuration is not a hint handed to an optimizer that is free to look elsewhere; it is the complete enumeration of what the run is allowed to consider, and anything outside it stays unreachable no matter how long the run is given. In a previous blog, Reverse-Engineering hipBLASLt TensileLite Kernels, we used that constraint defensively: decode a kernel you trust from its solution name, pin its roughly one hundred parameters as single-value ForkParameters, and the run cannot return anything slower, because the only kernel it can express is the one you started from. That gives tuning a floor.
Reverse-Engineering hipBLASLt TensileLite Kernels: From Solution Name to a Tuning Config
- 04 August 2026
In a previous blog, Customizing Kernels with hipBLASLt TensileLite GEMM Tuning, we showed how TensileLite Tuning generates brand-new GEMM kernels by searching a parameter space and selecting the fastest valid candidate for a given problem size. The search space is defined entirely by the tuning configuration YAML you provide. Its ForkParameters lists are expanded into a Cartesian product, and the tuning run can only ever return a kernel that is expressible as one combination within that product.
Building and Deploying Custom hipBLASLt Libraries on AMD Instinct GPUs
- 18 June 2026
General Matrix Multiply (GEMM) operations are a core component of many generative AI workloads. Whether you are running attention mechanisms in the prefill phase of a Large Language Model (LLM) or generating tokens sequentially during the decode phase, matrix multiplication performance has a direct impact on end-to-end latency and throughput.
Customizing Kernels with hipBLASLt TensileLite GEMM Tuning - Advanced User Guide
- 06 April 2026
Optimizing General Matrix Multiply (GEMM) operations is critical for maximizing the efficiency of AI models on AMD hardware. In our previous blog posts, we explored Offline Tuning, a method for selecting the best-performing kernel from an existing solution pool. For detailed instructions on using hipBLASLt-bench, please refer to hipBLASLt offline tuning part 1 and part 2. Additionally, for a streamlined experience, check out the Day 0 Developer Guide: hipBLASLt Offline GEMM Tuning Script which covers one-click offline tuning. Furthermore, for scenarios requiring dynamic runtime adaptation, developers can explore our recently published blog on hipBLASLt Online GEMM Tuning.
Debugging NaN Results in CK Tile GEMM: A rocgdb Detective Story
- 30 January 2026
When developing high-performance GPU kernels, subtle bugs can lead to catastrophic failures like NaN (Not-a-Number) outputs. This post chronicles our journey of debugging a tricky NaN issue in AMD’s Composable Kernel (CK) Tile GEMM implementation using rocgdb. What started as mysterious NaN outputs ended with discovering a single-character typo that corrupted the data distribution.
Avoiding LDS Bank Conflicts on AMD GPUs Using CK-Tile Framework
- 25 July 2025
LDS bank conflict is a common performance bottleneck in GPU kernel development. Composable Kernel (CK-Tile), a kernel development framework for AMD GPUs, provides a framework-level solution for LDS bank conflicts. Composable Kernel for ROCm is used to build portable high-performance kernels for accelerating computing, e.g. HPC, DL and LLMs for training and inference workloads. In this blog, we show you how to analyze, detect, and eliminate LDS bank conflicts using CK-Tile, AMD’s composable GPU kernel framework. A GEMM kernel serves as a classic example for analyzing how threads interact with LDS during both reads and writes. Starting with a naïve memory layout, we evaluate bank conflict behavior, explore mitigation techniques such as padding, and ultimately demonstrate how an XOR-based swizzle transformation achieves a bank conflict-free design.