GPT-6 reportedly nails SRE-Bench with ~100% pass@4 on never-public reverse-engineering binaries
xennygrimmato_ · x · 2026-09-04
SRE-Bench's author i2huer reveals OpenAI told them GPT-6 (Astra) essentially 'cooked' their software reverse-engineering benchmark: nearly 100% solve rate at pass@4 and 88% at pass@1, at lower cost. Since all benchmark programs are private, pre-2024 binaries, contamination is unlikely; the community is adding obfuscators and packers to harden it. The team is soliciting new private programs and anti-analysis tooling.
More from Models
- Deep Learning Weekly #471: Claude Fable 5.1 launch, production-parity LLM evals, alignment paper — dl_weekly · 2026-09-05
- OpenAI confirms Astra counts toward normal plan usage, users can allocate 100% of quota — soumitrashukla9 · 2026-09-05
- RareBench eval: Claude Fable 5.1 tops rare-disease diagnosis while Nemotron 3 Ultra scores 0% — danielmckinn0n · 2026-09-05
- GPT-6 Astra Beats 5.6 Sol Pro (Max) on FrontierMath T4; Open Models Seen 18 Months Behind — inductionheads · 2026-09-05
- TheZvi breaks down the Claude Fable 5.1 system card: 200+ pages of safety evals — TheZvi · 2026-09-05
- COLM paper: legible chain-of-thought steps aren't necessarily important — LauraRuis · 2026-09-05