Posts by Haishuo Kong
Hyperloom: A Multi-Agent Harness for Autonomous Inference Optimization on AMD GPUs
- 21 September 2026
ROCm™ Hyperloom delivered a median 1.73× inference speedup in extensive unattended evaluation on AMD Instinct™ GPUs, with gains ranging from 1.35× to 7.31×. It profiles each workload, searches framework and kernel optimizations, validates changes end to end, and carries proven results forward, without per-model human tuning.
Resilient Large-Scale Training: Integrating TorchFT with TorchTitan on AMD GPUs
- 08 February 2026
Training large AI models on AMD GPUs demands unwavering stability and robust fault-tolerance capabilities at cluster scale. Yet today’s ROCm-based multi-node GPU deployments often rely on brittle checkpoint-and-restart mechanisms to recover from failures. This approach wastes precious compute cycles and slows down training as model sizes and cluster scales grow. To address these challenges, we integrated PyTorch’s native fault-tolerance framework—TorchFT—with the TorchTitan training framework on AMD’s Primus-SaFE Kubernetes platform, achieving resilient, checkpoint-less training at hundred-GPU scale. This blog builds upon our previous work on the Primus ecosystem—for background on the platform architecture, see our earlier posts on Primus-SaFE, the Primus training framework, and training large models with Primus.