Astra scores 100% on ExploitBench while Fable 5.1 hits 55.8% on Terminal-Bench, but no head-to-head exists
heypearlai · x · 2026-09-06
From heypearlai's thread: Astra scores 100% on the cybersecurity benchmark ExploitBench, beating OpenAI's previous model's 78.5%, while Fable 5.1 scores 55.8% on Terminal-Bench, up from Fable 5's 42.0%. Caveat: the two have never been benchmarked against each other — each lab is grading its own homework.
More from Models
- New Astra model refuses to write election turnout-modeling code, citing 'predictive' concerns — sethlazar · 2026-09-06
- Grok Bot resets usage limits for all users over the long weekend — Baconbrix · 2026-09-06
- Rumor: Frontier Models May Be 48-Layer Transformers Looped Twice; DeepLoop Paper Explores Depth Scaling — dotey · 2026-09-06
- Ultra user: Astra ignores explicit instructions that Sol follows with the same prompt — PurpleManner5207 · 2026-09-06
- Frontier AI Is Now a Two-Company Race, Says Ethan Mollick — emollick · 2026-09-06
- 15 Minutes of CAD Work Decimates Claude Code's 5-Hour Usage Limit, User Finds — _Stocko_ · 2026-09-06