GPT-5 struggles enormously with batched moves in agent benchmarks, testers find
patience_cave · x · 2026-09-07
In an X thread, patiencecave reports that GPT-5 struggles enormously when allowed to submit batched moves in agent benchmarks — and it struggles even without batching.
Key points:
- FakePsyho, based on ARC-AGI-3 runs, argues batching still pays off: the cost savings on the 5.4-5.6 tier are massive despite quality dropping. Better to allow more steps plus batching than fewer steps with none.
- patiencecave notes batching makes sense on ARC-AGI-3 since a misstep is costly, and plans to retest on mazebench with one move per turn.
- The quality drop has so far been observed on Astra and Fable specifically; previous-generation models also perform badly when submitting several moves at once.
More from coding & agent
- LeanHttp: a libcurl-backed synchronous HTTP client for Lean 4 — hargup13 · 2026-09-07
- Emad Mostaque doubles down: AI will write nearly all code by 2027, developers gone by 2028 — jasonkneen · 2026-09-07
- Dingcad is dead: unsupervised Codex wrote the STEP export itself — yacineMTB · 2026-09-07
- Yacine confirms it's real: Codex CLI designed a welded chair and got a manufacturing quote — yacineMTB · 2026-09-07
- Bulletin: A Shared Message Board for Agents Over MCP and HTTP — MeAndClaudeMakeHeat · 2026-09-07
- Astra spins, guzzles tokens on large codebases; Fable 5.1 wins for hard coding — bindureddy · 2026-09-07