AI - Applications & Models#
Serving 64Mi-Token Contexts on One AMD Instinct™ MI355X Node
Explore how one 8-GPU AMD MI355X node serves Kimi Linear from 1K to 64M tokens under vLLM—and the TTFT and decode throughput behind the run.
Scaling RL with verl on AMD Instinct MI355X: Async Walkthrough and Sync Benchmark
Scale RLHF on AMD Instinct MI355X with verl's fully async trainer. Hands-on GRPO + DAPO examples.
Quark Support for HuggingFace Diffusers and SVDQuant
Learn how to quantize, save, and reload diffusion models in Quark using its new SVDQuant and HuggingFace Diffusers support.
VSA: Accelerating Video Diffusion Inference with Sparse Attention on AMD GPUs
Accelerate video diffusion inference with VSA sparse attention: up to 3.31x attention kernel-time speedup on AMD Instinct MI308X GPUs
Reverse-Engineering hipBLASLt TensileLite Kernels: From Solution Name to a Tuning Config
Pin the pool's best kernel into a TensileLite tuning config by decoding its solution name, so an expanded re-tune can only match or beat it.
Enabling Language-specific Reasoning in Multilingual Models with Reinforcement Learning
Learn how to train multilingual reasoning models with reinforcement learning and extend context windows on AMD Instinct GPUs.
Introducing Instella-MoE: A State-of-the-Art Fully Open Mixture-of-Experts Language Model
Explore Instella-MoE-16B-A3B, AMD’s fully open 16B MoE LLM with 2.8B active params per token, trained on AMD Instinct™ MI300 & MI325 GPUs.
Onboard and Deploy Custom Models in AMD AI Workbench
Learn how to deploy custom models in AMD AI Workbench, utilizing the AIM Engine, orchestration, scaling, profile parameters and API keys.
Efficient MiniMax-M3 Inference on AMD Instinct GPUs with ATOM and ATOMesh
Serve and benchmark MiniMax-M3 on AMD Instinct MI355X GPUs using ATOM and ATOMesh with EAGLE3 speculative decoding.
GEAK V3: Agent-Driven, Repository-Level GPU Kernel Optimization across HIP, Triton, and FlyDSL on AMD GPUs
Explore GEAK v3: agent-driven, repository-level GPU kernel optimization across HIP, Triton, and FlyDSL on AMD Instinct™ GPUs.
Multi-Accelerator Support for AIMs and AMD Solution Blueprints
Deploy and run AIMs and AMD Solution Blueprints across AMD Instinct™ GPUs, AMD EPYC™ CPUs, and AMD Radeon™ GPUs
Local Image and Video Generation on AMD Ryzen™ AI Max+ Processor (Windows)
Run ComfyUI natively on Windows on AMD Ryzen AI Max+ with ROCm 7.2.1—SDXL, Flux, and video workflows on the Radeon 8060S, no WSL.