Astra hits ≥98% on 10 of 37 benchmarks, near-saturating the time-horizon eval suite

dswg97 · x · 2026-09-09

The author describes their time-horizon (TH) estimation methodology: injecting synthetic long-horizon (>2 hour) tasks at 0% success rate, then filtering bootstrap samples with non-trivial success-rate variance for a more reasonable estimate.

Key caveat: Astra is close to saturating the suite — scoring ≥98% on 10 of 37 benchmarks, versus GPT-5.5 saturating only 4/37 — making a confident TH estimate for Astra hard to produce.

Related event: Astra nears benchmark saturation, scoring 98%+ on 10 of 37(2 posts)→

Original post →

More from Models

Models channel →