Miles now supports RL training for Qwen and GLM models with high-performance kernels

ying11231 · x · 2026-08-28

Miles now supports RL training for Alibaba's Qwen3.8-Flash-Next and Zai's GLM-5.3-Flash. By pairing SGLang rollout with Megatron training, it provides SGLang-consistent, high-performance kernels. It implements QSA indexing and sparse attention for Qwen, and KDA + DSA with kpool-compressed indexing for GLM, validated end-to-end on GB300 GPUs.

Original post →

More from Infra

Infra channel →