Posts by Zejun Chen

Scaling MiniMax-M3 Inference with Distributed Serving and Operator Co-Design on AMD Instinct MI355X GPUs

This blog walks you through a set of MiniMax-M3 inference optimizations on AMD Instinctâ„¢ MI355X GPUs using ATOM, AITER, and ATOMesh. For broader background on the inference engine and distributed serving layers used here, see the ROCm blogs on ATOM and ATOMesh.

Read more ...


vLLM-ATOM: Unlocking Native AMD Performance in the vLLM Ecosystem

This blog walks you through vLLM-ATOM, the AMD-optimized plugin that supercharges vLLM on Instinct GPUs.

Read more ...