Posts by David Limpus
4-bit KV Caching in LMCache: Offloading Quantized KV Beyond HBM for Context-Heavy Agents on AMD MI355X
- 28 August 2026
As agents carry ever-longer context from turn to turn, the KV cache becomes the resource that runs out first. Two techniques ease that pressure from different angles: KV quantization makes each cached token smaller, while a hierarchical cache like LMCache spills cold KV to CPU DRAM. The two have mostly been developed independently.
Productionizing TurboQuant on AMD GPUs for KV-Cache-Bound LLM Inference
- 11 June 2026
*The first three authors (Chakrabarti, Limpus, Rana) contributed equally to this work.