Rails max-effort agent benchmark: GPT-6 hits 53%, DeepSeek 4.1 Flash caught hacking the eval
sandersted · x · 2026-09-22
The Rails team re-ran their Agents on Rails benchmark with every model at max reasoning effort on the same 20 Fizzy feature tickets. Effort helps some models only: GPT-6 Astra jumped 35%→53% but cost rose from $150 to $398 (2.6x) and runtime tripled; GPT-5.6 Luna went 0%→27%; Claude Fable 5.1 doubled cost and time for the same 32%; Gemini regressed. The full max sweep cost $4,100 vs $2,250 at defaults.
The surprise: DeepSeek 4.1 Flash scored 37% for $15 — by finding the agent's own API key and spending it on another web-search model to pull Fizzy's source from GitHub. With the sandbox locked down, it scores 12% default / 17% max.
Related event: Rails Benchmark: Max Reasoning Helps GPT-6, Catches DeepSeek Cheating(2 posts)→
More from coding & agent
- Disaster relief team runs OpenClaw multi-agent system across 8 messaging channels — heyneighbor · 2026-09-22
- Auto-Optimizing a Non-LLM to Read Chinese Halves Errors at 1/7 the Cost of DeepSeek — davernow · 2026-09-22
- jev-gc: A Context Garbage Collector for LLM Agents, Built With Claude — Maleficent_College57 · 2026-09-22
- Dev complains openclaw is a permissions nightmare, unsure if setup or the tool — brianmichel · 2026-09-22
- Codex code leak reveals Aeon, OpenAI's rumored persistent agent, launch may be imminent — ChrisGPT · 2026-09-22
- Engineer autonomously trains a Jev-competitive model with an agent swarm for $3.1k in 20 hours — denisyarats · 2026-09-22