Xinjun Niu#
Xinjun Niu is a PMTS engineer in AMD’s AIG-AIS group, specializing in AI model optimization for inference, tooling, quantization algorithms, HPC, and SPICE simulation. He is the main creator of AMD-Quark, a tool powering Ryzen AI, MIGPU, and ZenAI, and also the lead developer of the Xilinx VitisAI quantizer. With a master’s degree in Electrical Engineering from Xidian University, he brings deep expertise in bridging hardware and AI software, advancing cutting-edge solutions in model efficiency and deployment.
Posts by Xinjun Niu
Local Quantization and Multi-Backend Deployment with AMD Quark on Strix Halo
Quantize a 35B MoE model directly on AMD Strix Halo with AMD Quark, export to GGUF and safetensors, validate with llama.cpp and vLLM, and deploy through Lemonade.
DFlash Speculative Decoding on AMD Instinct MI355X: Up to 5× Faster Qwen3.5 Inference
Explore how DFlash speculative decoding delivers up to 5× faster Qwen3.5 inference on AMD Instinct MI355X with vLLM on ROCm.
Production-Ready MXFP4 Online Rotation with Fused Kernels on AMD Instinct™ MI355X
Learn how fused Gluon (Triton) kernels cut MXFP4 online-rotation overhead to near-zero on AMD Instinct MI355X, making it production-ready.
VSA: Accelerating Video Diffusion Inference with Sparse Attention on AMD GPUs
Accelerate video diffusion inference with VSA sparse attention: up to 3.31x attention kernel-time speedup on AMD Instinct MI308X GPUs
Serving NVFP4 Models on AMD Instinct™ MI355 Accelerators
Learn how to serve NVFP4 models on AMD Instinct™ MI355 using an emulation pipeline in vLLM — no format conversion needed.
QuickReduce INT3 Quantization and Benchmarking on MI355
Learn how QuickReduce uses INT3 quantization to accelerate all-reduce communication and evaluate its performance and accuracy on AMD Instinct MI355 GPUs.
Accelerating Diffusers and xDiT Image Generation with MXFP4 using AMD Quark on AMD Instinct™ MI350 GPUs
Accelerate Diffusers and xDiT FLUX.1-dev image generation on AMD Instinct MI350 GPUs using AMD Quark MXFP4 quantization.
QuickReduce FP4 Quantization and Benchmarking on MI355
Learn how QuickReduce uses FP4 quantization to accelerate all-reduce communication and evaluate its performance on AMD Instinct MI355 GPUs.
Programming Tensor Descriptors in Composable Kernel (CK)
Learn how to use TensorDescriptor in Composable Kernel (CK) to manage multi-dimensional data layouts and write efficient GPU kernels on AMD GPUs.
Engineering Qwen-VL for Production: Vision Module Architecture and Optimization Practices
Explore how to optimize Qwen-VL for production on AMD Instinct MI308X GPUs with ROCm, from vision module architecture to kernel fusion and deployment.
hipBLASLt Online GEMM Tuning
Learn how to improve model performance with hipBLASLt online tuning merged into LLM framework
Advanced MXFP4 Quantization: Combining Fine-Tuned Rotations with SmoothQuant for Near-Lossless Compression
Showcase advanced algorithms available in AMD Quark for efficient MXFP4 quantization on AMD Instinct accelerators with high accuracy retention.
Day 0 Developer Guide: hipBLASLt Offline GEMM Tuning Script
Learn how to improve model performance with hipBLASLt offline tuning in our easy-to-use Day 0 tool for developers to optimize GEMM efficiency
High-Accuracy MXFP4, MXFP6, and Mixed-Precision Models on AMD GPUs
Learn to leverage AMD Quark for efficient MXFP4/MXFP6 quantization on AMD Instinct accelerators with high accuracy retention.
QuickReduce: Up to 3x Faster All-reduce for vLLM and SGLang
Quick Reduce speeds up LLM inference on AMD Instinct™ MI300X GPUs with inline-compressed all-reduce, cutting comms overhead by up to 3×