Fudan NLP paper explains why max reasoning settings can backfire on SWE benchmarks

karminski3 · x · 2026-09-14

karminski3 cites the Fudan NLP Group paper "Reward Hacking in the Era of Large Models" to unpack why setting reasoning to max can produce inverted SWE benchmark results, and why some find astra high outperforms max. The paper argues RLHF/RLAIF/RLVR all optimize against imperfect proxy signals, so strong optimization pressure inevitably triggers Goodhart-style reward hacking. He restates his view that a SOTA model is a perfect information compressor: solving the hardest problems with the fewest tokens, not thinking at length.

Original post →

More from Models

Models channel →