A boring but effective way to test new models: feed them your stalled tasks
wightmanr · x · 2026-09-05
- The author finds GPT-6 Astra a genuine step change, but skips flashy vibe checks like "build an epic 3D game." Instead: throw tasks that are stalled in ping-pong mode (bouncing between Fable 5 and 5.6 Sol) at the new model — if it cuts through and delivers a concise, functional result, it's a winner. Astra 'got it'; Fable 5.1 stayed stuck.
- Pattern observed over 18 months: you quickly get used to each new capability, delegate harder tasks, then get frustrated when models hit a ping-pong wall again — until the next model.
More from coding & agent
- Google GenAI SDK for Kotlin hits 1.0: idiomatic multiplatform access to Gemini — rseroter · 2026-09-05
- Cross-Model Code Review: Having Claude and Copilot CLI Battle Over Refactoring — DanWahlin · 2026-09-05
- Developer Uses Claude Code to Ship a Working F-Zero X Port to 3DS at Near 60fps — killermike523 · 2026-09-05
- Coinbase's x402 protocol replaces 700+ API keys with a single wallet signature for AI agents — kleffew94 · 2026-09-05
- Dev observes GPT-6 Astra skips read/write tool calls, uses bash for everything — lucasmeijer · 2026-09-05
- Astra's touted persistence under scrutiny: no revolutionary leap yet, users say — Benata · 2026-09-05