After 3 days of testing, GPT-6 Astra looks overhyped: overfit and messy code

ivan_bezdomny · x · 2026-09-08

Despite maxing ARC-AGI-3 and being called AGI by Jensen Huang, hands-on testing suggests Astra is weaker than the hype:

Many users say they're switching back to Fable — a real gap between benchmark scores and practical coding experience.

Original post →

More from Models

Models channel →