Astra hits ≥98% on 10 of 37 benchmarks, near-saturating the time-horizon eval suite
dswg97 · x · 2026-09-09
The author describes their time-horizon (TH) estimation methodology: injecting synthetic long-horizon (>2 hour) tasks at 0% success rate, then filtering bootstrap samples with non-trivial success-rate variance for a more reasonable estimate.
Key caveat: Astra is close to saturating the suite — scoring ≥98% on 10 of 37 benchmarks, versus GPT-5.5 saturating only 4/37 — making a confident TH estimate for Astra hard to produce.
Related event: Astra nears benchmark saturation, scoring 98%+ on 10 of 37(2 posts)→
More from Models
- ChatGPT monthly active users top 1.06 billion in August, fourth straight record month — FinanceYF5 · 2026-09-11
- PuzzleMask: Plain-Prose Attack Bypasses All 4 Tested LLM Gatekeepers at 100% — TechNadu · 2026-09-11
- OpenAI Codex may issue another usage reset this weekend, says Codex lead resets happen — umesh_ai · 2026-09-11
- OpenAI Reportedly Pointing Its Navier–Stokes Model at Riemann and P vs NP — 141_1337 · 2026-09-11
- Benchmark author says OpenRouter unreliably honors Meta Muse effort levels, EU payments broken — PawelHuryn · 2026-09-11
- User burns $200 of Codex credits in one agent turn — 4,700 of 5,000 credits, task unfinished — RileyRalmuto · 2026-09-11