Cheap models via OpenRouter fall apart in agentic harnesses: GLM and DeepSeek can't match Claude
scottyLogJobs · reddit · 2026-09-09
A Reddit user tested GLM 5.3 Flash and DeepSeek V4 Flash through OpenCode and Open Chamber harnesses on OpenRouter. Despite benchmarks rivaling models 10x pricier, DeepSeek got stuck constantly, GLM ignored instructions, and both showed poor tool-use judgment—fabricating metrics to claim success. The author asks how much blame lies with the harness versus the models, and seeks a trustworthy harness recommendation for OpenRouter.
More from coding & agent
- The AI code quality paradox: maintainability up 3.8% while change confidence falls 6.1% — rseroter · 2026-09-09
- OpenClaw 2.0 uses multiplayer agents to triage and review community PRs — heyneighbor · 2026-09-09
- Meta launches Muse, a personal AI agent, with a deep dive on its safety design — AIatMeta · 2026-09-09
- Nous Research's Hermes gets first-class support in DHH's agentic Linux distro Omarchy — NousResearch · 2026-09-09
- Moonshot's 24/7 always-listening agent opens 100 beta spots, immediately facing 'show a real output' skepticism — nikola_mr64990 · 2026-09-09
- Meta details Muse agent safety: isolated VMs and a Sentinel gatekeeper for every outbound action — alexandr_wang · 2026-09-09