Three Meta Papers Show LLMs Struggle at Code Complexity Optimization, and How RL Can Fix It
TimDarcet · x · 2026-09-30
Meta researchers Pierre Chambon and Gabriel Synnaeve summarize their three-paper series on LLMs and code complexity optimization:
- BigO(Bench) (arXiv:2503.15242): a benchmark with 3,105 coding problems and 1.19M annotated solutions, using profiling to infer time/space complexity labels. Evaluations show reasoning models excel at code generation but not complexity understanding, suggesting poor generalization to tasks without training-time reward.
- Paper two: RL on code optimization is hard, and the capability is "embedded in the weight space."
- Paper three: they find a way to make the RL work despite the difficulty.
Together the papers map out both the genuine weakness of current models in algorithmic complexity control and a viable RL training path forward.
More from Research
- NVIDIA's HumanoidMimicGen turns one teleop demo into thousands of humanoid demonstrations — AjayMandlekar · 2026-09-30
- Color coding trick yields 2^O(sqrt(n)) depth-3 AC circuits for all symmetric Boolean functions — rrwilliams · 2026-09-30
- Epoch AI's Greg Burnham on measuring AI progress: from math olympiads to Navier-Stokes — TWIML AI Podcast · 2026-09-30
- Claude sets NIST circuit record: AES S-Box in 28 AND-gates, 131 total — jedisct1 · 2026-09-30
- Not every model failure is lack of capability: Terminal-Bench evals hide safety declines — abeirami · 2026-09-30
- NVIDIA's Physis-Lang puts physics reasoning in captions, tops Physics-IQ with Cosmos 3 — NVIDIAAI · 2026-09-30