GPT-6 Astra tops ZeroBench, surpassing human baseline across all three metrics
Waiting4AniHaremFDVR · reddit · 2026-09-23
GPT-6 Astra makes a massive leap on ZeroBench, an extremely difficult vision benchmark, surpassing the human baseline on all three metrics. The poster clarifies the scoring: pass@5 counts if at least one of 5 attempts is correct; pass^5 requires all 5 correct (a reliability metric); and the reported "pass@1" is actually the average score across 5 attempts, not a true single-attempt pass@1.
More from Models
- Rumor: Opus 5.5 Said to Be Cheaper and Faster Than Astra — BLUECOW009 · 2026-09-23
- The industry dropped the ball on strong, tool-calling, non-reasoning small LLMs — vboykis · 2026-09-23
- Asked AI to Review a Known Concurrency Bug, It Came Back Clean — cto_junior · 2026-09-23
- Developer Stress-Tests Opus 5.5 With Dozens of Tasks Before It Gets Nerfed — remilouf · 2026-09-23
- Jev API explodes at $0.042/M tokens: a hands-on checklist from desktop agents to drone control — blaizedsouza · 2026-09-23
- swyx Makes Opus 5.5 the Default for AINews After Head-to-Head Test, Industry Cuts Prices 40-50% — rickasaurus · 2026-09-23