Adil Lashab#
I work on LLM inference on Instinct GPUs (MI Family). Most of it is running real models at scale and fixing what breaks or runs slow: speculative decoding, MoE, FP8, attention/MLA, KV cache.
Posts by Adil Lashab
I work on LLM inference on Instinct GPUs (MI Family). Most of it is running real models at scale and fixing what breaks or runs slow: speculative decoding, MoE, FP8, attention/MLA, KV cache.