Jiejing Zhang#
Jiejing Zhang has extensive experience in building large-scale AI inference platforms and leading low-level LLM inference optimization. His expertise spans high-performance computing, GPU architecture, distributed inference, KV-cache systems, and production-scale model serving. He currently serves as an AI inference architect at AMD, focusing on inference acceleration and high-performance serving frameworks.
Posts by Jiejing Zhang