Applications & models#
Explore the latest blogs about applications and models in the ROCm ecosystem, including machine learning frameworks, AI models, and application case studies.
Reproducing AMD MLPerf Inference v6.1 Submission Results
In this blog, we share the technical details of how we accomplish the results in our MLPerf Inference v6.1 submission.
Technical Dive into AMD MLPerf Inference v6.1 Submission
Learn about the ROCm optimizations powering dlrm-v3, llama2-70b, and gpt-oss-120b performance.
DFlash Speculative Decoding on AMD Instinct MI355X: Up to 5× Faster Qwen3.5 Inference
Explore how DFlash speculative decoding delivers up to 5× faster Qwen3.5 inference on AMD Instinct MI355X with vLLM on ROCm.
Knowledge Graph Integration With Poro2 For Enriching Medical Text Processing
Learn how to integrate a medical knowledge graph with Poro2 LLM via MCP to simplify medical text on AMD Instinct MI300X GPUs
Serving 64Mi-Token Contexts on One AMD Instinct™ MI355X Node
Explore how one 8-GPU AMD MI355X node serves Kimi Linear from 1K to 64M tokens under vLLM—and the TTFT and decode throughput behind the run.
Scaling RL with verl on AMD Instinct MI355X: Async Walkthrough and Sync Benchmark
Scale RLHF on AMD Instinct MI355X with verl's fully async trainer. Hands-on GRPO + DAPO examples.
Quark Support for HuggingFace Diffusers and SVDQuant
Learn how to quantize, save, and reload diffusion models in Quark using its new SVDQuant and HuggingFace Diffusers support.
VSA: Accelerating Video Diffusion Inference with Sparse Attention on AMD GPUs
Accelerate video diffusion inference with VSA sparse attention: up to 3.31x attention kernel-time speedup on AMD Instinct MI308X GPUs
Reverse-Engineering hipBLASLt TensileLite Kernels: From Solution Name to a Tuning Config
Pin the pool's best kernel into a TensileLite tuning config by decoding its solution name, so an expanded re-tune can only match or beat it.
Enabling Language-specific Reasoning in Multilingual Models with Reinforcement Learning
Learn how to train multilingual reasoning models with reinforcement learning and extend context windows on AMD Instinct GPUs.
Introducing Instella-MoE: A State-of-the-Art Fully Open Mixture-of-Experts Language Model
Explore Instella-MoE-16B-A3B, AMD’s fully open 16B MoE LLM with 2.8B active params per token, trained on AMD Instinct™ MI300 & MI325 GPUs.
Onboard and Deploy Custom Models in AMD AI Workbench
Learn how to deploy custom models in AMD AI Workbench, utilizing the AIM Engine, orchestration, scaling, profile parameters and API keys.