OpenAI's Astra Scored 62.7% and 99.9% on the Same Benchmark, 37 Points Apart
mixtapedmonk · reddit · 2026-09-14
Digging into ARC Prize's actual results table, the author found OpenAI's GPT-6 Astra posted 62.7% and 99.9% on the same ARC-AGI-3 benchmark depending on which harness was used—37 points apart—and the org that built the test won't call it AGI. Fortune separately found five numbers quietly changed on OpenAI's own launch page after it went live. Comparing the Llama 4 benchmark precedent, the writeup concludes it could be genuine harness noise or something else, but headline-screenshot benchmark reporting is clearly broken.
More from Models
- ZDTaichu5.0-9B, a 9B vision-language model with spatial reasoning, trends on Hugging Face — TaichuAI · 2026-09-15
- Atria Dawn Preview: student-heavy team launches research-focused agentic base model — xiaohu · 2026-09-15
- Which 10Eros video-model quant works best on 8GB VRAM? A practical trade-off question — apostrophefee · 2026-09-15
- rasbt shows why final-result benchmarks mislead: Astra vs Qwen in Paint — rasbt · 2026-09-15
- Cristóbal Valenzuela praises Solaris: 'Websites are going to be fun again' — c_valenzuelab · 2026-09-15
- Google DeepMind on speech-to-speech: conversational, intelligent, multimodal — pick two — AI Engineer · 2026-09-15