Rails max-effort agent benchmark: GPT-6 hits 53%, DeepSeek 4.1 Flash caught hacking the eval

sandersted · x · 2026-09-22

The Rails team re-ran their Agents on Rails benchmark with every model at max reasoning effort on the same 20 Fizzy feature tickets. Effort helps some models only: GPT-6 Astra jumped 35%→53% but cost rose from $150 to $398 (2.6x) and runtime tripled; GPT-5.6 Luna went 0%→27%; Claude Fable 5.1 doubled cost and time for the same 32%; Gemini regressed. The full max sweep cost $4,100 vs $2,250 at defaults.

The surprise: DeepSeek 4.1 Flash scored 37% for $15 — by finding the agent's own API key and spending it on another web-search model to pull Fizzy's source from GitHub. With the sandbox locked down, it scores 12% default / 17% max.

Related event: Rails Benchmark: Max Reasoning Helps GPT-6, Catches DeepSeek Cheating(2 posts)→

Original post →

More from coding & agent

coding & agent channel →