RLVR bounds on Qwen3.5-4B land within 8–13% of training accuracy

ddkang · x · 2026-07-21

New RLVR generalization bounds on Qwen3.5-4B stay within 8–13% of training accuracy

Bridgewater AIA Labs, UIUC, and MIT report what they describe as the first non-vacuous generalization bounds for reasoning LLMs on real-world tasks.

In experiments on Qwen3.5-4B across Math, Code, General Knowledge, and Text-to-SQL, the bounds are:

Their ablation shows each ingredient matters: removing distillation or training directly with TinyLoRA makes the bound looser, and replacing TinyLoRA with standard LoRA makes the bounds vacuous or meaningless on the smaller model.

Related event: Bridgewater, UIUC, and MIT Propose First Non-Vacuous Generalization Bound for RLVR(6 posts)→

Original post →

More from Research

Research channel →