Posts by Akash Haridas

Zebra-HyLo: Upcycling Transformers into Long-Context Hybrid LLMs on AMD Instinctâ„¢ GPUs

Hybrid language models that interleave attention with linear sequence-modeling blocks have become the default answer to the cost of long context. Jamba, Samba, Qwen3-Next and Kimi-Linear all take this shape, and they all share one property that makes them expensive to adopt: they are pretrained from scratch. Every one of them pays the full cost of building a foundation model again, which means the enormous investment already sunk into existing Transformer checkpoints is thrown away.

Read more ...


Nitro-T: Training a Text-to-Image Diffusion Model from Scratch in 1 Day

AMD is excited to release Nitro-T, a family of text-to-image diffusion models focused on highly efficient training. Our models achieve competitive scores on image generation benchmarks compared to previous models focused on efficient training while requiring less than 1 day of training from scratch on 32 AMD Instinct MI300X GPUs.

Read more ...