Posts by Huasha Zhao

Technical Dive into AMD MLPerf Inference v6.1 Submission

MLPerf Inference v6.1 results were released on September 16, 2026. For AMD, this was a highly successful round in which we achieved several leadership scores, expanded benchmark coverage, introduced a new GPU with validated performance, and enabled record-setting results from our partners. In this round, AMD and its partners provided validated results across a diverse set of workloads, including dlrm-v3 (Deep Learning Recommendation Model v3), llama2-70b (Meta’s 70-billion-parameter language model), gpt-oss-120b (an open-weight 120-billion-parameter Mixture-of-Experts model), deepseek-r1 (DeepseekAI’s 671-billion-parameter Mixture-of-Experts model), and Wan 2.2-t2v (a 14-billion-parameter text-to-video generation model). Single-node results used 8 GPUs per node; gpt-oss-120b was additionally submitted at multi-node scale across 72 AMD Instinct™ MI355X GPUs. All results were validated through MLPerf peer review.

Read more ...


Reproducing AMD MLPerf Inference v6.1 Submission Results

This blog shows you how to reproduce AMD submission results for MLPerf Inference v6.1 on AMD Instinct MI355X, MI350X, and MI350P GPUs using self-contained Docker images, publicly available quantized model weights, and a step-by-step benchmark recipe.

Read more ...


Dropless MoE Training in JAX with Primus-Turbo

Mixture-of-Experts (MoE) models have become a standard way to scale a transformer’s parameter count without paying the full compute bill — but training them efficiently on GPUs forces an uncomfortable trade-off. The default path in JAX/MaxText keeps every expert’s tensors at a fixed shape and simply drops the tokens that overflow each expert’s capacity, trading model quality for speed. The fully dropless alternative keeps every token, but in pure JAX it hits a memory wall that makes it impractical at production scale.

Read more ...


MaxText-Slurm: Production-Grade LLM Training with Built-In Observability

Training large language models (LLMs) at scale on GPU clusters is not just a compute problem — it is an operations problem. Launching multi-node distributed training, keeping it running reliably, and diagnosing failures when they happen all require tooling that most training frameworks do not provide. MaxText-Slurm is an open-source launch system and observability stack that bridges this gap for MaxText on AMD Instinct GPU clusters managed by Slurm.

Read more ...