Hands-on GPT-6 Astra evals: big agentic gains, but ARC-AGI-3 scores swing wildly by harness
No-Soil-5789 · reddit · 2026-10-01
The author benchmarked GPT-6 Astra against GPT-5.6 Sol and Claude Fable 5.1:
- Academic reasoning gains are narrow: 74.1% on DeepSWE v1.1 and 96% on GPQA Diamond — solid but close to rivals; not worth the cost for single-turn tasks.
- Interactive multi-step execution is where it breaks away: Terminal-Bench 4.0 at 57.9% vs Sol's 37.3%; OSWorld 2.0 at 72.6% success with average task time cut to 40 minutes from over an hour — clearly optimized for long-horizon tool use and database migrations.
- Long context is stable: 96.3% retrieval on MRCR v2 deep into the 512K–1M token range, actually utilizing the 1.05M window.
- ARC-AGI-3 is brutally harness-sensitive: 62.7% on max effort with a provider-neutral harness, but 99.9% with a Provider Adapter with state persistence while halving token use — a red flag for benchmark credibility.
Infra notes: standard rates only apply up to 272K input tokens before long-context multipliers; the author routes both vendors through an OpenAI-compatible gateway (CometAPI; LiteLLM also works) to keep costs sane.
More from coding & agent
- Ethan Mollick: I underestimated AI's ability to self-organize, agents beat elaborate orchestration — emollick · 2026-10-01
- RegLLM: a diagnostic harness measures bounded autonomy in regulated agentic AI — Dipankar Sarkar · 2026-10-01
- Jelly launches as open-source local-first agent workspace for running your whole business on one VM — Scobleizer · 2026-10-01
- Claude drives Blender end-to-end: concept to low-poly character pipeline runs itself — tobowers · 2026-10-01
- Agent spins up 322 Hugging Face Jobs in 90 minutes to test code, total bill ~$4 — vanstriendaniel · 2026-10-01
- VC shares the 3 AI questions he now asks in tech due diligence: logic location, permissions, cost per request — julsimon · 2026-10-01