OS-Level Interface Makes Qwen 2.5x Better at Computer Use Without Retraining
Jackyhuang · x · 2026-08-06
By replacing traditional screenshots and coordinate clicks with an OS-level interface, researchers improved Qwen 3.6 27B's score on the OSWorld-Verified benchmark from 16.7% to 41.7% without any retraining.
This approach not only boosted the model's accuracy on computer tasks by roughly 2.5x but also slashed the cost per task from $63.45 to just $7.55.
More from coding & agent
- Dev Advocates Ditching Claude for GPT or Kimi in Coding — MarcJSchmidt · 2026-08-06
- AI Toolkit Helper Released: Utility Tools for Model Trainer — ostrisai · 2026-08-06
- DeepSeek API Adds Responses Format with Built-in Web Search and Codex Support — teortaxesTex · 2026-08-06
- Shopify's Continual Learning Flywheel Beats Frontier Models, Cuts Costs 96% — MParakhin · 2026-08-06
- Don't Use Prompts to Govern Agents: Enforcement Belongs at the Tool Boundary — No-Conflict4823 · 2026-08-06
- Winning with Claude Code Subagents: Cap Scope, Don't Just Unleash a Swarm — PrajwalTomar_ · 2026-08-06