Posts by Huasha Zhao
Technical Dive into AMD MLPerf Inference v6.1 Submission
- 17 September 2026
MLPerf Inference v6.1 results were released on September 16, 2026. For AMD, this was a highly successful round in which we achieved several leadership scores, expanded benchmark coverage, introduced a new GPU with validated performance, and enabled record-setting results from our partners. In this round, AMD and its partners provided validated results across a diverse set of workloads, including dlrm-v3 (Deep Learning Recommendation Model v3), llama2-70b (Meta’s 70-billion-parameter language model), gpt-oss-120b (an open-weight 120-billion-parameter Mixture-of-Experts model), deepseek-r1 (DeepseekAI’s 671-billion-parameter Mixture-of-Experts model), and Wan 2.2-t2v (a 14-billion-parameter text-to-video generation model). Single-node results used 8 GPUs per node; gpt-oss-120b was additionally submitted at multi-node scale across 72 AMD Instinct™ MI355X GPUs. All results were validated through MLPerf peer review.
Reproducing AMD MLPerf Inference v6.1 Submission Results
- 17 September 2026
This blog shows you how to reproduce AMD submission results for MLPerf Inference v6.1 on AMD Instinct MI355X, MI350X, and MI350P GPUs using self-contained Docker images, publicly available quantized model weights, and a step-by-step benchmark recipe.
Dropless MoE Training in JAX with Primus-Turbo
- 10 June 2026
Mixture-of-Experts (MoE) models have become a standard way to scale a transformer’s parameter count without paying the full compute bill — but training them efficiently on GPUs forces an uncomfortable trade-off. The default path in JAX/MaxText keeps every expert’s tensors at a fixed shape and simply drops the tokens that overflow each expert’s capacity, trading model quality for speed. The fully dropless alternative keeps every token, but in pure JAX it hits a memory wall that makes it impractical at production scale.
MaxText-Slurm: Production-Grade LLM Training with Built-In Observability
- 02 March 2026
Training large language models (LLMs) at scale on GPU clusters is not just a compute problem — it is an operations problem. Launching multi-node distributed training, keeping it running reliably, and diagnosing failures when they happen all require tooling that most training frameworks do not provide. MaxText-Slurm is an open-source launch system and observability stack that bridges this gap for MaxText on AMD Instinct GPU clusters managed by Slurm.