Posts by Yu Shao

Benchmarking Kimi-K3 Across vLLM, SGLang, and ATOM on MI350X

Moonshot AI released the weights for Kimi-K3, a 2.8-trillion-parameter, 1M-context, natively-MXFP4 Mixture-of-Experts (MoE) model, and AMD Instinct™ support was available on day 0 across three independent serving frameworks: vLLM, SGLang, and ATOM. The recipes target the gfx950 generation (MI350X and MI355X). The measurements in this post were taken on an 8× MI350X node.

Read more ...