Local models on one RTX 5090 now match Claude Code on real-task agent benchmark
dh7net · reddit · 2026-10-05
A new benchmark (airbench.ai) tests model+harness combos on math, vision, computer use, and coding, measuring both capability and speed.
Key findings:
- Claude Code + Opus 5.5 scores 100% (14m26s); local zcode/4xRTX6k/glm-5.3-flash-NVFP4 also hits 100% (31m13s)
- Speed pick qwen3.8-flash-next-iq3xxs-strata: 96% in 7m41s — faster than Claude Code
- On local hardware, the harness matters as much as the model: same quant on the same 5090 scores 22%–96% depending on harness
- Best harnesses: opencode (94% avg) and omp (91%); hermes is often too slow; MTP variants and 65k context hurt scores
- Best single-5090 models: swift-1.5-qwen3.8-27b-q6k (most robust, 82% avg across 5 harnesses) and qwen3.8-flash-next-strata
Anyone can contribute configs via a generated prompt.
More from coding & agent
- Daniel Miessler Launches Deeds to Replace PRs as the Unit of AI Coding Work — karlwaldman · 2026-10-05
- Pi 1.0 Ships with Durable Harness: Agent Frameworks Turn Into Distributed Systems — ghumare64 · 2026-10-05
- Vibe Coders Won: Everyone Now Just Talks to Agents in a Chat Box — bindureddy · 2026-10-05
- 10 AI Agent Projects That Take You From Tool Calling to Autonomous Systems — _jaydeepkarale · 2026-10-05
- 7 Claude Prompts to Learn Anything Faster, From Planning to Flashcards — CodeByPoonam · 2026-10-05
- Pi adds MCP support after publicly mocking it—engineering team explains what changed — bibryam · 2026-10-05