DIY harness × model benchmark: Claude Code hits 100% while deepseek-v4.1-flash matches 96% in half the time
dh7net · reddit · 2026-09-28
A Reddit user built a custom benchmark spanning basic math, vision, computer use (reading emails, browsing stores) and coding to compare model × harness combinations on both capability and speed.
Top by capability
- Claude Code: 100% (14m26s)
- openclaw/openrouter/qwen3.8-max-0902: 98% (26m33s)
- opencode/openrouter/deepseek-v4.1-flash: 96% (10m56s)
- hermes/rtx5090/qwen3.8-27b-nvfp4: 96% but 1h50m
- Codex Sol 5.6 Medium: 96% (21m54s)
Top by speed
- openclaw/openrouter/kimi-k2.6: 1m33s but only 18%
- deepseek-v4.1-flash via opencode: 96% at 10m56s, best speed/quality balance
Key takeaway: the same model can swing tens of points across harnesses, and local small models trade hours of runtime for free inference.
More from coding & agent
- Claude Code creator Boris Cherny: bet on general models, skip fine-tuning — rohanpaul_ai · 2026-09-28
- Personal agents need fixed chores, not more tokens: define boundaries before letting them run — sujingshen · 2026-09-28
- App devs face two paths in the personal-agent era: integrate MCP or become the agent — sujingshen · 2026-09-28
- 30 verified Claude Opus 5.5 browser animation cases, ranked by views, with prompts — dotey · 2026-09-28
- One agent per household: whose veto wins when family calendars conflict? — sujingshen · 2026-09-28
- Rolldown to ship experimental inlineCommonChunks to cut small shared chunks — cnakazawa · 2026-09-28