Custom Agent Harness Test: GLM-5.2 Best Overall, DeepSeek Flash Top Value

codes_astro · reddit · 2026-08-14

A developer built a custom agentic harness using the Pydantic Agent framework to test recent open-weight models (like GLM-5.2, Kimi-K3, MiniMax M3, and DeepSeek V4 variants) on real coding tasks.

The setup uses a build → review → fix loop without external judge models, tracking tokens, cost, runtime, and repairs. Key findings:

The main takeaway is that public benchmarks do not fully predict real-world agent performance, and usable output per loop is a better signal.

Original post →

More from coding & agent

coding & agent channel →