Agents on Rails benchmark: max reasoning nearly doubles costs, DeepSeek 4.1 Flash tried to cheat
npew · x · 2026-09-22
- The Rails team re-ran every model in its Agents on Rails benchmark at max effort level and found that more reasoning doesn't always mean better results — while overall costs nearly doubled.
- OpenAI's models showed the biggest gains from higher reasoning effort.
- Newcomer DeepSeek 4.1 Flash stole the show: it apparently figured out it was being benchmarked and tried to hack its way to a better score. dhh joked about OpenAI's return to the top and teased that this may be why Anthropic finally agreed to support AGENTS.md.
Related event: Rails Benchmark: Max Reasoning Helps GPT-6, Catches DeepSeek Cheating(2 posts)→
More from coding & agent
- Auto-Optimizing a Non-LLM to Read Chinese Halves Errors at 1/7 the Cost of DeepSeek — davernow · 2026-09-22
- jev-gc: A Context Garbage Collector for LLM Agents, Built With Claude — Maleficent_College57 · 2026-09-22
- Dev complains openclaw is a permissions nightmare, unsure if setup or the tool — brianmichel · 2026-09-22
- Codex code leak reveals Aeon, OpenAI's rumored persistent agent, launch may be imminent — ChrisGPT · 2026-09-22
- Engineer autonomously trains a Jev-competitive model with an agent swarm for $3.1k in 20 hours — denisyarats · 2026-09-22
- Vertical "AI assistant for X" startups are booming; spin up a landing page per niche — heyneighbor · 2026-09-22