TerminalBench scoring isn't comparable: Astra gets zeros on safety stops, Fable re-routes to Opus
xeophon · x · 2026-09-15
xeophon flagged an inconsistency in public TerminalBench traces: when Astra hits safety classifiers, its rollouts are stopped and scored as 0, whereas Fable gets re-routed to Opus and the run continues. Such differing handling makes leaderboard scores hard to compare fairly.
More from Models
- Leaker teases 'more exciting OpenAI releases this week' at DevDay-level ship volume — MickeySteamboat · 2026-09-15
- Mistral audio research lead: voice AI needs a screen, won't replace it — Machine Learning Street Talk · 2026-09-15
- User argues uneven AI subscription tiers: $200 plan gets 20× usage vs 5× at $100 — NoOne_n13 · 2026-09-15
- OpenAI reportedly pauses $200 ChatGPT Pro signups as GPT-6 Astra demand saturates capacity — emmanuelvivier · 2026-09-15
- Vibe coding 3D game dev: Claude Pro beats ChatGPT Plus on limits and code audits — AdvertisingBubbly546 · 2026-09-15
- Quant Finance Shows What a Scaling-Pilled AI Industry Looks Like — and How the Moat Fades — willcb · 2026-09-15