Hands-on GPT-6 Astra evals: big agentic gains, but ARC-AGI-3 scores swing wildly by harness

No-Soil-5789 · reddit · 2026-10-01

The author benchmarked GPT-6 Astra against GPT-5.6 Sol and Claude Fable 5.1:

Infra notes: standard rates only apply up to 272K input tokens before long-context multipliers; the author routes both vendors through an OpenAI-compatible gateway (CometAPI; LiteLLM also works) to keep costs sane.

Original post →

More from coding & agent

coding & agent channel →