Posts by Yuankai Chen
Debugging Logprob Mismatches in LLM Reinforcement Learning
- 23 September 2026
In online reinforcement learning (RL), a language model generates responses, receives rewards, and updates its weights to improve future responses. Many systems split this loop between a rollout engine, which generates the responses, and a trainer, which learns from them. Both engines evaluate token probabilities: rollout records the log-probability, or logprob, of each token it generates, while the trainer evaluates those same tokens when computing its loss. After the update, the new weights are sent back to rollout, and the loop begins again.
Implementing a High-Performance Custom Diffusion Attention Kernel with FlyDSL
- 17 September 2026
Readers may be familiar with traditional Transformer models and their attention mechanisms. The traditional autoregressive transformers generate tokens iteratively. Since this feature significantly limits the inference throughput, researchers have begun exploring approaches such as diffusion models that can generate multiple tokens in each iteration.
veRL on AMD: Production-Ready RL Post-Training on ROCm
- 08 September 2026
Reinforcement learning post-training on AMD Instinct GPUs is here — with a turnkey container, AITER-accelerated vLLM and SGLang rollout, and accuracy validated on both MI300 and MI355.
Scaling RL with verl on AMD Instinct MI355X: Async Walkthrough and Sync Benchmark
- 18 August 2026
Reinforcement learning (RL) for large language models (LLMs) alternates between two phases: generation (rollout), where the current policy produces responses, and training, where those responses are used to update the policy. In verl, the key design choices are when these phases run relative to each other (synchronously or with overlap) and where they run (colocated on the same GPUs or on separate GPU pools). This blog first explains the differences between the two modes and when to use each.
Primus Tuning Agent: Closing the Configuration-Search Loop
- 06 July 2026
Error parsing meta tag attribute “keywords”: No content.
Primus Projection: Estimate Memory and Performance Before You Train
- 24 April 2026
Error parsing meta tag attribute “keywords”: No content.
Primus-Pipeline: A More Flexible and Scalable Pipeline Parallelism Implementation
- 23 February 2026
Error parsing meta tag attribute “keywords”: No content.
MoE Training Best Practices on AMD GPUs
- 16 December 2025
This blog covers best practices for training Mixture-of-Experts (MoE) models on AMD Instinct™ MI300/MI355-series[a] GPUs with the ROCm ecosystem. Whether you’re new to MoE distributed architectures or optimizing trillion-parameter models, this guide will help you identify bottlenecks and maximize efficiency on AMD hardware.