Astra+Codex beats human testers on cost: $5.6 per puzzle game, 60x cheaper than verified runs
andreisavu · x · 2026-09-06
Researcher @FakePsyho ran Astra + Codex on a public sandboxed demo (no internet) and tallied the actual cost:
- All tools available: $140 total ≈ $5.6/game, 60-70x cheaper than the verified run
- No code provided: $240 ≈ $9.6/game, 35-40x cheaper
- Compare with ARC-AGI-3 at $10-15/game and $35-40 per completion
The author argues AI is now the first model more cost-effective than humans at puzzle games: a full verified run that would cost $20K with the original setup should run $500-1K via Codex, i.e. $10-20 per game — roughly matching human tester pay. Takeaway: stop calling AI stupid and move to harder benchmarks.
More from Models
- Astra's release triggered a breakout weekend of AI-CAD: from rebuilding SF landmarks to multi-tool 3D workflows — burhop · 2026-09-06
- Dan Jeffries: embedded reasoning will kill the CoT-distillation accusations against Chinese models — Dan_Jeffries1 · 2026-09-06
- Alexandr Wang urges users to try Muse Spark 1.3 max, called 'Opus-level' by users — alexandr_wang · 2026-09-06
- Models would rather edit files with raw Python than use edit tools — xeophon · 2026-09-06
- ChatGPT agent can click through 'I'm not a robot' checks, reigniting AI security debate — GaelVaroquaux · 2026-09-06
- Astra's language beats Sol by a mile: less QA jargon, docs worth actually sharing — raskingballs · 2026-09-06