Jiejing Zhang

Jiejing Zhang

Jiejing Zhang#

Jiejing Zhang has extensive experience in building large-scale AI inference platforms and leading low-level LLM inference optimization. His expertise spans high-performance computing, GPU architecture, distributed inference, KV-cache systems, and production-scale model serving. He currently serves as an AI inference architect at AMD, focusing on inference acceleration and high-performance serving frameworks.

Posts by Jiejing Zhang

https://rocm.blogs.amd.com/software-tools-optimization/infera-di/README.html