Gemini 3.6 Flash Beats GPT-5.6 in Browser Agent Benchmark
allenainie · x · 2026-07-22
According to browseruse's benchmark tests, Google's newly released Gemini 3.6 Flash achieves a score of 68% in web agent tasks. It outperforms GPT-5.6-sol (67%) and Sonnet 4.6 (62%), ranking second only to Opus 4.8 (74%) at a fraction of the cost, offering frontier-level browser agents at Flash pricing.
More from coding & agent
- Atomic Agent targets local open-model use with 1,000+ skills and MCP servers — eyishazyer · 2026-07-23
- video-use edits raw footage into final.mp4 by reading transcripts, not frames — alex_verem · 2026-07-23
- Citrolabs opens ego-lite, a browser for humans and AI agents to work in parallel — citrolabs · 2026-07-23
- Alibaba open-sources open-code-review, a hybrid LLM agent for line-level code review — alibaba · 2026-07-23
- Codex feels cleaner on a blank project than inside a messy legacy codebase — Individual-Carob5593 · 2026-07-23
- LoHoSearch turns a 7.62M-entity knowledge graph into a harder benchmark for search agents — 美团技术团队 · 2026-07-23