Posts by Hang Yang

DFlash Speculative Decoding on AMD Instinct MI355X: Up to 5× Faster Qwen3.5 Inference

Autoregressive decode generates one token per model forward pass, so single-request latency is bounded by how fast you can stream weights through the GPU — not by compute. Speculative decoding attacks that bottleneck by letting a small draft model propose several tokens that the large target model then verifies in a single parallel pass. In this post we bring DFlash [3] — a block-diffusion drafter — to AMD Instinct MI355X through vLLM [4] on ROCm [5], benchmark it head-to-head against Qwen3.5’s [6] built-in MTP drafter, and then quantize the target model to mxfp4 to show that speculation and quantization stack.

Read more ...


Programming Tensor Descriptors in Composable Kernel (CK)

Writing efficient GPU kernels requires more than knowing the API—it demands a deep understanding of the underlying concepts, from GPU architecture to low-level programming patterns. This blog series demystifies GPU kernel programming on AMD GPUs by breaking down common kernels into their fundamental building blocks. Rather than treating GPU programming as a black box, each blog focuses on a specific concept, starting from first principles and building up to complete implementations with simple, insightful example code. In this blog, you will learn one of the most fundamental concepts in Composable Kernel (CK): the TensorDescriptor—a powerful abstraction for managing multi-dimensional data layouts and transformations. By the end of this series, you will be able to not only understand existing GPU kernels but also design and optimize your own.

Read more ...