Posts by Aswin Mathews
Benchmarking Kimi-K3 Across vLLM, SGLang, and ATOM on MI350X
- 22 September 2026
Moonshot AI released the weights for Kimi-K3, a 2.8-trillion-parameter, 1M-context, natively-MXFP4 Mixture-of-Experts (MoE) model, and AMD Instinct™ support was available on day 0 across three independent serving frameworks: vLLM, SGLang, and ATOM. The recipes target the gfx950 generation (MI350X and MI355X). The measurements in this post were taken on an 8× MI350X node.
DI Series: Scaling GLM-5.1-FP8 to 64 MI300X GPUs
- 21 August 2026
Serving a frontier Mixture-of-Experts (MoE) model well is a systems problem, and it gets harder the moment one node is not enough. GLM-5.1 is a good example: it is a large, sparse MoE that users want to run at long context, and it ships a new attention family that breaks assumptions older serving stacks quietly relied on. Fitting it on eight GPUs is only the start. The real question is how to keep it correct and fast as you spread it across several nodes.