LMSYS Debuts Miles: Blackwell-Native 8-bit and 4-bit RL Recipes

BanghuaZ · x · 2026-07-30

LMSYS has introduced two Blackwell-native, low-precision reinforcement learning (RL) recipes within the Miles framework, fully open-sourced in collaboration with @humansand and @nvidia.

Key technical highlights include:

Testing on Qwen3-30B-A3B shows that all 5 low-precision configurations closely track the BF16 reward curve while significantly reducing rollout time.

Original post →

More from Infra

Infra channel →