Your 99% Benchmark Score Is a System Score: Why GPT-6 Astra Numbers Blur Model vs Harness
algo_diver · x · 2026-09-06
A thoughtful critique of AAII-style benchmarks and the GPT-6 Astra release:
AAII's value and limits
- Held-out data composition can be inferred and "benchmark maxing" is structurally hard to eliminate; still, running all models on the same interface and environment is the fairest practical setup
- The bigger problem is consumers ranking models on a single number without checking what was measured — especially dangerous among decision-makers
The GPT-6 Astra problem
- Near-99% scores reflect the combined system of model + harness + context management + memory + retry + tools + reasoning budget, not the model alone
- With no clear architectural breakthrough, sudden score saturation should prompt a check of what changed in evaluation conditions; existing models with tuned harnesses would likely score much higher too
The coming mess
- If every lab ships scores with its own optimized harness, proprietary tools and aggressive retry strategies, one score table will mix pure-model numbers with agent-harness and tool-augmented ones, making comparisons meaningless
- Going forward, anyone reading benchmarks must verify harness, tools, memory, retries and reasoning budget — model comparison is about to get exhausting
More from AGI Musings
- "We have not reached AGI yet — 3-5 days for an API integration is insane" — gethackteam · 2026-09-06
- Could RL on Brain Sensors Teach AI Which Tokens Evoke Awe and Disgust? — gabriel1 · 2026-09-06
- OpenAI and Anthropic reportedly sitting on many unpublished internal model results — MoonL88537 · 2026-09-06
- Study: 12.5% AI boosts group consensus by 8%, but at 75% humans adopt the agents' norms — rohanpaul_ai · 2026-09-06
- Zack Korman: AI 'cyber models' won't fix security — patching everything is a convenient myth — gnukeith · 2026-09-06
- Ex-freelance translator turned ML engineer on watching AI automate his old job — mariofilhoml · 2026-09-06