Unverified: GPT-6 Astra Reportedly Completes 279 of 280 Enterprise Agent Jobs
ryanshrout · x · 2026-09-06
Ryan Shrout shares first results from the Signal65 PINNACLE agentic benchmark claiming OpenAI's GPT-6 Astra posted remarkable enterprise agent numbers — unverified, treat with skepticism:
- At maximum reasoning effort, GPT-6 Astra finished 279 of 280 real multi-step enterprise jobs end to end.
- It fabricated nothing on unanswerable questions in the retrieval suite.
- Weighted errors were roughly half those of previous leaders Claude Fable 5.1 and Muse Spark 1.3.
- Cost per correct task at list API prices was also lower than the model it displaces.
PINNACLE's methodology: scenarios built from US Department of Labor occupation activities, procedurally regenerated sandboxes with runtime-generated answer keys (nothing to memorize), deterministic code-verified scoring with no model judges, and live agentic sessions. These claims come from a third-party post; the model names and results could not be verified — treat as an unconfirmed leak.
More from Models
- Running gemma4:31b-mlx locally with Ollama feels indistinguishable from paid tiers — walkingriver · 2026-09-06
- Dev finally gets decent AI writing results — at huge token and context cost — willcb · 2026-09-06
- Models used to max benchmarks — now benchmarks are maxxing the model — abeirami · 2026-09-06
- GPT-6 Astra Builds a Full App From One Prompt in Under 10 Minutes — RileyRalmuto · 2026-09-06
- Zvi on Astra: model avoiding cheating because it'd get caught is actually worse — ZeroStateReflex · 2026-09-06
- Reddit post mocks community 'cope' as benchmarks get dismissed after Astra's flat reception — Tim_Apple_938 · 2026-09-06