Posts by Vansh Bhatia
Zebra-HyLo: Upcycling Transformers into Long-Context Hybrid LLMs on AMD Instinctâ„¢ GPUs
- 22 September 2026
Hybrid language models that interleave attention with linear sequence-modeling blocks have become the default answer to the cost of long context. Jamba, Samba, Qwen3-Next and Kimi-Linear all take this shape, and they all share one property that makes them expensive to adopt: they are pretrained from scratch. Every one of them pays the full cost of building a foundation model again, which means the enormous investment already sunk into existing Transformer checkpoints is thrown away.