Astra scores 77.3% on Browser Use Benchmark v2, crushing Opus 5's 50.5%
gabrielchua · x · 2026-09-06
Browser Use Benchmark v2 results show Astra (medium) at 77.3%, far ahead of Opus 5 at 50.5% and GPT-5.6 Sol xhigh at 49.1%. 22 of 60 tasks earned full marks, while Opus 5 got zero perfect scores. The author notes 'fable' was excluded for rejecting too many tasks.
More from Models
- Ollama CEO: open models will carry 80-90% of enterprise tokens at just 10-20% of cost — victor_explore · 2026-09-06
- Users report Chat model update brought obsessive over-complication loops — marcoshsq · 2026-09-06
- GPT Desktop Users Spot New 'Astra' Model in Drop-Down, Praise X-High Fast Mode — beffjezos · 2026-09-06
- Dev burns 20B tokens in a week on personal account, praises crazy Codex usage resets — rudrank · 2026-09-06
- Dev's Take on Model Releases: Mild General Gains Plus Specialized Data Tricks — xiaosun86 · 2026-09-06
- Astra shows striking gains on visual-spatial intelligence benchmark SpatialBench — socoolandawesome · 2026-09-06