AI Blogs#
Iteratively Tuning hipBLASLt TensileLite Kernels: A Smaller Search, a Faster Kernel
Learn why a small iterative search beats a big one-shot sweep for hipBLASLt TensileLite kernels, in less time and with no risk of regression.
veRL on AMD: Production-Ready RL Post-Training on ROCm
Run veRL RL post-training on AMD Instinct GPUs with a turnkey ROCm container, AITER-accelerated rollout, and validated accuracy on MI300 and MI355.
Efficiently Serving NVFP4 Models on AMD Instinct™ MI350X/MI355X Accelerators via Online NVFP4 to Quark MXFP4 Requantization
Serve NVFP4 models on MI350X/MI355X via SGLang's online NVFP4 to MXFP4 requantization: no preprocessing, minimal accuracy impact, native throughput.
Enabling DeepSeek-V4-Flash Training on AMD Instinct MI355X GPUs with Primus
DeepSeek-V4-Flash training on AMD Instinct GPUs with Primus: model architecture introduction, performance projection, kernel optimizations, and how to reproduce.
Serving 64Mi-Token Contexts on One AMD Instinct™ MI355X Node
Explore how one 8-GPU AMD MI355X node serves Kimi Linear from 1K to 64M tokens under vLLM—and the TTFT and decode throughput behind the run.
Scaling RL with verl on AMD Instinct MI355X: Async Walkthrough and Sync Benchmark
Scale RLHF on AMD Instinct MI355X with verl's fully async trainer. Hands-on GRPO + DAPO examples.
Quark Support for HuggingFace Diffusers and SVDQuant
Learn how to quantize, save, and reload diffusion models in Quark using its new SVDQuant and HuggingFace Diffusers support.
VSA: Accelerating Video Diffusion Inference with Sparse Attention on AMD GPUs
Accelerate video diffusion inference with VSA sparse attention: up to 3.31x attention kernel-time speedup on AMD Instinct MI308X GPUs
Optimizing ATOM and vLLM-ATOM for High-Interactivity Inference
Cut per-token latency on AMD Instinct MI355X GPUs: see how ATOM, vLLM-ATOM, and AITER strip fixed costs out of the LLM decode path.
4-bit KV Caching in LMCache: Offloading Quantized KV Beyond HBM for Context-Heavy Agents on AMD MI355X
Combining 4-bit KV quantization (TurboQuant) with KV offload (LMCache) on AMD MI355X nearly doubles serving goodput for long-context agents.
DI Series: Scaling GLM-5.1-FP8 to 64 MI300X GPUs
Scale GLM-5.1-FP8 long-context serving beyond a single node on AMD Instinct MI300X with WideEP prefill-decode disaggregation and MoRI.
Exploring XGBoost: A Deep Dive
A single source explanation of XGBoost features and its working principles.
AUP Learning Cloud: Streamlining AI Education on AMD
An all-in-one ROCm JupyterHub platform that deploys GPU-ready AI teaching environments on AMD hardware with one installer and open-source teaching labs.
Introducing AMD CDNA™ 5 and the AMD Helios™ Rackscale Solution
Introducing AMD CDNA 5, the AMD Instinct MI455X GPU, and the AMD Helios rackscale solution: an open, integrated platform for rack-scale AI.
Styled Text Image Generation with Eruku on AMD
Hands-on, reproducible guide to train and run Eruku on LUMI supercomputer, powered by AMD Instinct MI250X GPUs.
Elevate Your LLM Inference: Autoscaling with Ray, ROCm 7.0.0, and SkyPilot
Learn how to use multi-node and multi-cluster autoscaling in the Ray framework on ROCm 7.0.0 with SkyPilot
Stay informed
- Subscribe to our RSS feed (Requires an RSS reader available as browser plugins.)
- Signup for the ROCm newsletter
- View our blog statistics