Fudan NLP paper explains why max reasoning settings can backfire on SWE benchmarks
karminski3 · x · 2026-09-14
karminski3 cites the Fudan NLP Group paper "Reward Hacking in the Era of Large Models" to unpack why setting reasoning to max can produce inverted SWE benchmark results, and why some find astra high outperforms max. The paper argues RLHF/RLAIF/RLVR all optimize against imperfect proxy signals, so strong optimization pressure inevitably triggers Goodhart-style reward hacking. He restates his view that a SOTA model is a perfect information compressor: solving the hardest problems with the fewest tokens, not thinking at length.
More from Models
- David Bellamy clarifies his experiment used K2 Horizon, an open-weights 375B LLM — JeremyNguyenPhD · 2026-09-14
- Swift-Qwen3.8-27b, a token-efficient reasoning Qwen finetune, trends on Hugging Face — ukisai · 2026-09-14
- GPT-6 Astra hands-on: composes first, orchestrates later, and reportedly outshines Fable and Sol — paw_lean · 2026-09-14
- Researcher posts proof he both synthesized viruses and trained a 375B open-weight LLM — ethanCaballero · 2026-09-14
- Toby Ord: 10x more RLVR compute cuts tokens-to-target ~3x; gains may be math-specific — tobyordoxford · 2026-09-14
- New scaling curve has half the slope: 10,000x compute for 20%-to-80%, but bigger generational jumps — tobyordoxford · 2026-09-14