Astra scores 100% on ExploitBench, prompting a harder internal variant benchmark
inductionheads · x · 2026-09-03
Andrew Ginns reports that Astra scored a perfect 100% on ExploitBench, a benchmark for exploit capabilities, prompting the team to build an internal variant to better measure the model's capabilities—implying the original benchmark has saturated.
More from Models
- Ethan Mollick tests Gemini 3.8 Flash: fast but no match for Fable 5.1 on shaders — eldonredwards · 2026-09-03
- Meta's Spark 1.3 nears frontier and its scorched-earth pricing may threaten AI labs, argues investor — RihardJarc · 2026-09-03
- Claude Fable 5.1 (high) hits 92.3% on WeirdML, beating Fable 5 by 0.4% for new SOTA — teortaxesTex · 2026-09-03
- Would OpenAI bet $100M+ on Looped Transformer for Astra without scaling proof? — teortaxesTex · 2026-09-03
- Enterprises pay 10x-20x more to keep data out of AI training, suggesting a routing-layer startup — random_walker · 2026-09-03
- antirez: coding benchmarks are 'mostly crap' and badly misaligned with how models train — antirez · 2026-09-03