Astra scores 77% vs Fable 5.1's 56% on browser use, refusals blamed for gap
steipete · x · 2026-09-08
Sharing a browser-use benchmark comparison: Astra scores 77%, well ahead of Fable 5.1 at 56%. The key differentiator is that Fable "just refuses too much" — excessive task refusals drag down its completion rate. A reminder that in agentic browser tasks, willingness to act matters as much as raw capability.
Related event: Astra Tops Browser Use Benchmark at 77%(2 posts)→
More from Models
- ChatGPT monthly active users top 1.06 billion in August, fourth straight record month — FinanceYF5 · 2026-09-11
- PuzzleMask: Plain-Prose Attack Bypasses All 4 Tested LLM Gatekeepers at 100% — TechNadu · 2026-09-11
- OpenAI Codex may issue another usage reset this weekend, says Codex lead resets happen — umesh_ai · 2026-09-11
- OpenAI Reportedly Pointing Its Navier–Stokes Model at Riemann and P vs NP — 141_1337 · 2026-09-11
- Benchmark author says OpenRouter unreliably honors Meta Muse effort levels, EU payments broken — PawelHuryn · 2026-09-11
- User burns $200 of Codex credits in one agent turn — 4,700 of 5,000 credits, task unfinished — RileyRalmuto · 2026-09-11