Recent Posts - Page 2#
Local Quantization and Multi-Backend Deployment with AMD Quark on Strix Halo
Quantize a 35B MoE model directly on AMD Strix Halo with AMD Quark, export to GGUF and safetensors, validate with llama.cpp and vLLM, and deploy through Lemonade.
Debugging Logprob Mismatches in LLM Reinforcement Learning
Learn to debug logprob mismatches between rollout and training with Qwen3 examples on AMD Instinct MI355X GPUs.
Performance Profiling on AMD GPUs - Part 6: Advanced Thread Trace (ATT) - The Microscope for Your Application
Explore rocprofv3 ATT and ROCprof Compute Viewer to trace GPU kernels and explain stalls, waits, and memory-bound performance.
Zebra-HyLo: Upcycling Transformers into Long-Context Hybrid LLMs on AMD Instinct™ GPUs
Upcycle pretrained Transformers into long-context hybrid MLA + linear models on AMD Instinct MI300X GPUs, with 14 open checkpoints and training code.
Benchmarking Kimi-K3 Across vLLM, SGLang, and ATOM on MI350X
Learn how one declarative madengine command benchmarks day-0 Kimi-K3 across vLLM, SGLang, and ATOM on AMD Instinct MI350X.
Hyperloom: A Multi-Agent Harness for Autonomous Inference Optimization on AMD GPUs
Hyperloom is a multi-agent harness that autonomously optimizes LLM inference on AMD Instinct GPUs, reaching a median 1.73x throughput gain.
Serving GLM-5.2-MXFP4 on AMD Instinct™ MI355X: When Prefill Context Parallelism Pays
Learn when prefill context parallelism pays on AMD Instinct MI355X: 43-54% more throughput on long prompts with GLM-5.2-MXFP4, and when it loses.
Implementing a High-Performance Custom Diffusion Attention Kernel with FlyDSL
Learn how to implement and optimize flexible, high-performance diffusion attention kernels with FlyDSL.
Reproducing AMD MLPerf Inference v6.1 Submission Results
In this blog, we share the technical details of how we accomplish the results in our MLPerf Inference v6.1 submission.
Technical Dive into AMD MLPerf Inference v6.1 Submission
Learn about the ROCm optimizations powering dlrm-v3, llama2-70b, and gpt-oss-120b performance.
DFlash Speculative Decoding on AMD Instinct MI355X: Up to 5× Faster Qwen3.5 Inference
Explore how DFlash speculative decoding delivers up to 5× faster Qwen3.5 inference on AMD Instinct MI355X with vLLM on ROCm.
Thread Trace Part 1: ROCprof Compute Viewer
Learn to capture thread traces with rocprofv3 and analyze instruction timing, stalls, utilization, and counters in ROCprof Compute Viewer.