Featured Posts
Technical Dive into AMD MLPerf Inference v6.1 Submission
Learn about the ROCm optimizations powering dlrm-v3, llama2-70b, and gpt-oss-120b performance.
A Deep Dive into LDS Optimizations on AMD Instinct MI450 GPUs
Learn how to optimize LDS traffic in Gluon kernels on AMD Instinct MI450 GPUs using transposed loads and partition-conflict-free layouts.
ROCm 10.0: A Decade of Open Compute, Built for the Age of Agentic AI
Explore what's new in ROCm 10.0, headlined by ROCm.AI, the ROCm CLI, AMD Skills, and Hyperloom, alongside the ROCm Core SDK and platform-wide improvements
ROCm 7.14: TheRock Goes Production and Expands AMD's AI Software Platform
Explore what's new in ROCm 7.14: TheRock goes production, expanded hardware support, stronger AI frameworks, and enhanced profiling tools.
Implementing a High-Performance Custom Diffusion Attention Kernel with FlyDSL
Learn how to implement and optimize flexible, high-performance diffusion attention kernels with FlyDSL.
Reproducing AMD MLPerf Inference v6.1 Submission Results
In this blog, we share the technical details of how we accomplish the results in our MLPerf Inference v6.1 submission.
DFlash Speculative Decoding on AMD Instinct MI355X: Up to 5× Faster Qwen3.5 Inference
Explore how DFlash speculative decoding delivers up to 5× faster Qwen3.5 inference on AMD Instinct MI355X with vLLM on ROCm.
Thread Trace Part 1: ROCprof Compute Viewer
Learn to capture thread traces with rocprofv3 and analyze instruction timing, stalls, utilization, and counters in ROCprof Compute Viewer.
Knowledge Graph Integration With Poro2 For Enriching Medical Text Processing
Learn how to integrate a medical knowledge graph with Poro2 LLM via MCP to simplify medical text on AMD Instinct MI300X GPUs
Serving 64Mi-Token Contexts on One AMD Instinct™ MI355X Node
Explore how one 8-GPU AMD MI355X node serves Kimi Linear from 1K to 64M tokens under vLLM—and the TTFT and decode throughput behind the run.
Scaling RL with verl on AMD Instinct MI355X: Async Walkthrough and Sync Benchmark
Scale RLHF on AMD Instinct MI355X with verl's fully async trainer. Hands-on GRPO + DAPO examples.
Quark Support for HuggingFace Diffusers and SVDQuant
Learn how to quantize, save, and reload diffusion models in Quark using its new SVDQuant and HuggingFace Diffusers support.
An Educational GEMM Ladder for Helios GPUs
Build high-performance BF16 GEMM kernels on Helios GPUs with HipKittens, from a naive baseline to optimized schedules.
Iteratively Tuning hipBLASLt TensileLite Kernels: A Smaller Search, a Faster Kernel
Learn why a small iterative search beats a big one-shot sweep for hipBLASLt TensileLite kernels, in less time and with no risk of regression.
veRL on AMD: Production-Ready RL Post-Training on ROCm
Run veRL RL post-training on AMD Instinct GPUs with a turnkey ROCm container, AITER-accelerated rollout, and validated accuracy on MI300 and MI355.
Efficiently Serving NVFP4 Models on AMD Instinct™ MI350X/MI355X Accelerators via Online NVFP4 to Quark MXFP4 Requantization
Serve NVFP4 models on MI350X/MI355X via SGLang's online NVFP4 to MXFP4 requantization: no preprocessing, minimal accuracy impact, native throughput.
Enabling Physical AI Agents with Lemonade
Learn to deploy local agents for interactive robot arm manipulation using the Lemonade framework.
AUP Learning Cloud: Streamlining AI Education on AMD
An all-in-one ROCm JupyterHub platform that deploys GPU-ready AI teaching environments on AMD hardware with one installer and open-source teaching labs.
Introducing AMD CDNA™ 5 and the AMD Helios™ Rackscale Solution
Introducing AMD CDNA 5, the AMD Instinct MI455X GPU, and the AMD Helios rackscale solution: an open, integrated platform for rack-scale AI.
ROCm 7.13: Expanding Hardware, Tools, and Reach
Explore what's new in the ROCm 7.13 release, featuring expanded hardware support, GPU virtualization, enhanced developer tooling, and TheRock's modular packaging.
Stay informed
- Subscribe to our RSS feed (Requires an RSS reader available as browser plugins.)
- Signup for the ROCm newsletter
- View our blog statistics
- View the ROCm Developer Hub
- Report an issue or request a feature
- We are eager to learn from our community! If you would like to contribute to the ROCm Blogs, please submit your technical blog for review at our GitHub. Blog creation can be started through our GitHub user guide.