Simular's Agent Sai Tops OSWorld 2.0, Beating GPT-5.6 at Two-Thirds the Cost
xwang_lk · x · 2026-08-28
Simular's computer agent Sai achieved a 73% success rate on OSWorld 2.0, ahead of GPT-5.6 Sol (62.57%, per OpenAI) and Opus 5 (70.57%, per Anthropic), at roughly two-thirds the cost.
OSWorld 2.0, released in June by HKU's XLANG Lab, has 108 long-horizon tasks taking skilled humans over an hour each: cross-source reasoning (receipts scattered across email and expense reports), responding to dynamic changes, following tutorials precisely, and troubleshooting contradictory data.
Sai operates desktop apps and webpages, calls APIs, and writes code like a human — orchestrating a mixture of frontier and specialist models with dedicated perception/action interfaces, running on Simular's remote fleet or personal machines.
Related event: Simular's Sai Agent Tops OSWorld 2.0 at Two-Thirds the Cost(4 posts)→
More from coding & agent
- Practitioner take: cross-product MCP workflows are missing; agents could run 60-80% of company ops — pritisinghhhh · 2026-08-28
- Dev newsletter: 'My AI assistant is turning into an operating system' — dSebastien · 2026-08-28
- Grok Bot Tested: Reuses Hermes Patterns, Controls Any API — gregmushen · 2026-08-28
- iris launches: parallel agents edit your site in-browser and push PRs — floguo · 2026-08-28
- First device wallet for AI agents enables autonomous crypto signing — Daniel_Farinax · 2026-08-28
- Claude Code Safety Fails: Attack Succeeds 80% and Blocks Cleanup — Simon Willison · 2026-08-28