Developers Blogs#
Debugging Logprob Mismatches in LLM Reinforcement Learning
Learn to debug logprob mismatches between rollout and training with Qwen3 examples on AMD Instinct MI355X GPUs.
Performance Profiling on AMD GPUs - Part 6: Advanced Thread Trace (ATT) - The Microscope for Your Application
Explore rocprofv3 ATT and ROCprof Compute Viewer to trace GPU kernels and explain stalls, waits, and memory-bound performance.
Hyperloom: A Multi-Agent Harness for Autonomous Inference Optimization on AMD GPUs
Hyperloom is a multi-agent harness that autonomously optimizes LLM inference on AMD Instinct GPUs, reaching a median 1.73x throughput gain.
Thread Trace Part 1: ROCprof Compute Viewer
Learn to capture thread traces with rocprofv3 and analyze instruction timing, stalls, utilization, and counters in ROCprof Compute Viewer.
Iteratively Tuning hipBLASLt TensileLite Kernels: A Smaller Search, a Faster Kernel
Learn why a small iterative search beats a big one-shot sweep for hipBLASLt TensileLite kernels, in less time and with no risk of regression.
Enabling DeepSeek-V4-Flash Training on AMD Instinct MI355X GPUs with Primus
DeepSeek-V4-Flash training on AMD Instinct GPUs with Primus: model architecture introduction, performance projection, kernel optimizations, and how to reproduce.
Optimizing ATOM and vLLM-ATOM for High-Interactivity Inference
Cut per-token latency on AMD Instinct MI355X GPUs: see how ATOM, vLLM-ATOM, and AITER strip fixed costs out of the LLM decode path.
ROCm 10.0: A Decade of Open Compute, Built for the Age of Agentic AI
Explore what's new in ROCm 10.0, headlined by ROCm.AI, the ROCm CLI, AMD Skills, and Hyperloom, alongside the ROCm Core SDK and platform-wide improvements
Exploring XGBoost: A Deep Dive
A single source explanation of XGBoost features and its working principles.
Memory Instruction Scheduling for Lock-Stepped Kernels on AMD Instinct™ MI300X: Introducing the Series
Explore how instruction scheduling eases stalls on AMD Instinct MI300X in a new blog series. This intro covers the motivating example and methodology.
Bring Claude Code On‑Prem with AMD Instinct GPUs
Start running Claude Code securely with a self-hosted SGLang LLM on AMD Instinct MI355X GPUs.
AUP Learning Cloud: Streamlining AI Education on AMD
An all-in-one ROCm JupyterHub platform that deploys GPU-ready AI teaching environments on AMD hardware with one installer and open-source teaching labs.