LMSYS Debuts Miles: Blackwell-Native 8-bit and 4-bit RL Recipes
BanghuaZ · x · 2026-07-30
LMSYS has introduced two Blackwell-native, low-precision reinforcement learning (RL) recipes within the Miles framework, fully open-sourced in collaboration with @humansand and @nvidia.
Key technical highlights include:
- End-to-end MXFP8 across rollout, forward, and backward GEMMs.
- Hardware-native NVFP4 W4A4 RL specifically optimized for MoE models.
- Bit-exact quantizers between training and rollout to minimize mismatch.
- Fine-grained, cross-stack per-layer precision control.
Testing on Qwen3-30B-A3B shows that all 5 low-precision configurations closely track the BF16 reward curve while significantly reducing rollout time.
More from Infra
- AI Infrastructure Spending Outpaces Cash Flow: Google's Capex Up 107% — Beth_Kindig · 2026-07-30
- Cerebras on the Agentic Era: New Workflows Will Drive Non-GPU Chip Architectures — sarahookr · 2026-07-30
- Cognition Lab Talk: RL and Inference Optimization Are Converging — AAAzzam · 2026-07-30
- Vector Institute Demystifies MoE: Slashes Logit Memory from 23.3GB to 0.3GB — VectorInst · 2026-07-30
- Deploying LTX Video Models on Cloud GPUs: Pitfalls and an Automated Installer — Humble_Cut6799 · 2026-07-30
- NVIDIA Expected to Raise GeForce RTX GPU Prices Again by Up to 30% — ANR2ME · 2026-07-30