Meta paper says RL can optimize code speed, with Qwen 2.5 7B and CWM 32B gains

burny_tech · x · 2026-07-29

Meta paper shows how to make reinforcement learning optimize code speed

This arXiv paper from Meta FAIR looks at a hard problem: once you add execution time to the reward, RL for code optimization becomes brittle because timing noise, sparse rewards, and GRPO instability overwhelm the signal.

What they changed

Reported results

The core takeaway is that optimizing for speed is possible, but only if the test harness, reward shaping, and optimizer are all engineered for noisy execution-time signals.

Original post →

More from Models

Models channel →