Software tools & optimizations#
Discover the latest blogs about ROCm software tools, libraries, and performance optimizations to help you get the most out of your AMD hardware.
4-bit KV Caching in LMCache: Offloading Quantized KV Beyond HBM for Context-Heavy Agents on AMD MI355X
Combining 4-bit KV quantization (TurboQuant) with KV offload (LMCache) on AMD MI355X nearly doubles serving goodput for long-context agents.
A Deep Dive into LDS Optimizations on AMD Instinct MI450 GPUs
Learn how to optimize LDS traffic in Gluon kernels on AMD Instinct MI450 GPUs using transposed loads and partition-conflict-free layouts.
DI Series: Scaling GLM-5.1-FP8 to 64 MI300X GPUs
Scale GLM-5.1-FP8 long-context serving beyond a single node on AMD Instinct MI300X with WideEP prefill-decode disaggregation and MoRI.
Exploring XGBoost: A Deep Dive
A single source explanation of XGBoost features and its working principles.
Memory Instruction Scheduling for Lock-Stepped Kernels on AMD Instinct™ MI300X: Introducing the Series
Explore how instruction scheduling eases stalls on AMD Instinct MI300X in a new blog series. This intro covers the motivating example and methodology.
Bring Claude Code On‑Prem with AMD Instinct GPUs
Start running Claude Code securely with a self-hosted SGLang LLM on AMD Instinct MI355X GPUs.
Production-Ready MXFP4 Online Rotation with Fused Kernels on AMD Instinct™ MI355X
Learn how fused Gluon (Triton) kernels cut MXFP4 online-rotation overhead to near-zero on AMD Instinct MI355X, making it production-ready.
Using ODC to Accelerate AMD SFT Training
Learn how we ported ODC on-demand P2P communication to AMD Instinct MI300X to cut FSDP bubbles and speed up variable-length SFT training.
Closing the GPU Cluster Validation Gap: A Kubernetes-Native Approach with CVF
Learn how to validate AMD GPU clusters end-to-end with CVF: hardware acceptance, mesh bandwidth, RDMA, and RCCL testing in one pipeline.
AMD GPU Operator v1.5.0: DRA Support, Automated GPU Node Recovery, and Expanded Kubernetes Infrastructure Control
Discover how AMD GPU Operator v1.5.0 improves GPU scheduling, automates node recovery, and expands Kubernetes control.
Attention Decode on AMD MI450 GPUs: A Gluon Kernel Optimization Guide
Learn how to design a high-performance attention decode kernel on AMD MI450 GPUs using Gluon.
Serve Kimi-K2.5-MXFP4 on MI355X with ATOM
Serve Kimi-K2.5-MXFP4 on MI355X with ATOM and gfx950 block-scaled FP4 kernels for optimized LLM inference.