Posts by Aref Jafari
Zebra-HyLo: Upcycling Transformers into Long-Context Hybrid LLMs on AMD Instinct™ GPUs
- 22 September 2026
Hybrid language models that interleave attention with linear sequence-modeling blocks have become the default answer to the cost of long context. Jamba, Samba, Qwen3-Next and Kimi-Linear all take this shape, and they all share one property that makes them expensive to adopt: they are pretrained from scratch. Every one of them pays the full cost of building a foundation model again, which means the enormous investment already sunk into existing Transformer checkpoints is thrown away.
Serving 64Mi-Token Contexts on One AMD Instinct™ MI355X Node
- 24 August 2026
Error parsing meta tag attribute “keywords”: No content.