Browsecomp Score Correction: Sol Single-Agent Hits 90.4
nrehiew_ · x · 2026-07-17
Correction: Sol is actually SOTA on Browsecomp.
- Single-agent score is 90.4.
- Multi-agent score is 92.2.
The focus of this update isn't a new method, but an accurate correction of previous results, clarifying the score differences between the single-agent and multi-agent setups.
More from coding & agent
- Harness engineering is emerging as the execution layer for reliable AI agents — Pavan_Belagatti · 2026-07-21
- DevFest Lisbon keynote will cover Google AI Studio’s latest vibe coding and agentic AI features — gerardsans · 2026-07-21
- Daniel Hanchen’s 2-hour workshop covers open models, reward hacking and RL — danielhanchen · 2026-07-21
- DocETL adds a Python DSL and new tools to clean up long-form LLM outputs — sh_reya · 2026-07-21
- A Codex joke turns into a recursive debate about Cloud Codex — Dimillian · 2026-07-21
- AI-assisted development is widening security backlogs, so teams should measure risk velocity — WeldPond · 2026-07-21