Claude burns 24h building cells that take seconds to measure, in overnight benchmark fail
StefanoGogioso · x · 2026-09-04
Stefano Gogioso shares a hilarious overnight agent fail: when asked why the benchmarks weren't done, Claude's own diagnosis was damning — at a 2^24 grid size, each cell took 1-5 hours to build but under 1 second to measure. A 24-hour run yielded only a few hundred microsecond timings.
Beyond the comedy, it's a cautionary tale about letting models autonomously design experiments without any cost-awareness of imbalanced build-vs-measure budgets.
More from coding & agent
- Microsoft Foundry Model Router: one deployment that picks the best LLM per request in real time — lee_stott · 2026-09-04
- One-prompt game builds: Astra ships open-world adventure with branching dialogue in browser — eyishazyer · 2026-09-04
- Same-prompt race: Astra builds GTA VI-style game in 90 minutes vs Fable 5.1's 2 hours — eyishazyer · 2026-09-04
- Hermes cuts agent codebase context cost by ~63% after Teknium's cleanup pass — Teknium · 2026-09-04
- Claude Fable 5 extended to July 12 — 3 high-leverage ways to use it before access ends — femke_plantinga · 2026-09-04
- Celigo says AI agents built almost all of Ora, with humans approving every merge — Crescitaly · 2026-09-04