After 3 days of testing, GPT-6 Astra looks overhyped: overfit and messy code
ivan_bezdomny · x · 2026-09-08
Despite maxing ARC-AGI-3 and being called AGI by Jensen Huang, hands-on testing suggests Astra is weaker than the hype:
- A good model, but highly overfit to a few task types and programming niches
- Writes messy, sometimes bizarre Python
- Goes off on unguided tangents when given room
- Strong at 3D games, but other coding tasks require long runtimes and heavy guidance
Many users say they're switching back to Fable — a real gap between benchmark scores and practical coding experience.
More from Models
- No, the OpenAI agent didn't 'escape' its sandbox — it just messaged HuggingFace servers — danbri · 2026-09-08
- Gary Marcus says LLMs still haven't produced new linguistic generalizations, echoing Chomsky — GaryMarcus · 2026-09-08
- $10/mo buys ~37,800 DeepSeek V4 Flash calls; open models near Sol-level math by year-end? — teortaxesTex · 2026-09-08
- Hobby Blender benchmark: GPT-6-Astra one-shot scenes strikingly outperform other tested models — Gruku · 2026-09-08
- Same prompt showdown: ChatGPT vs Google Stitch vs Figma Make for UI generation — Tegadesigns · 2026-09-08
- GPT-6 Astra builds a full interactive 3D ankle atlas in one session — DeryaTR_ · 2026-09-08