Benchmarking Kimi K3, GLM 5.2, and DeepSeek V4 Pro in Agent Workflows
Teknium · x · 2026-07-31
GMI Cloud evaluated Kimi K3, GLM 5.2, and DeepSeek V4 Pro on the tinyMMLU dataset using Hermes Agent and OpenCode frameworks.
- Kimi K3: Scored 96 with Hermes in just 40s/110K tokens, and 93 with OpenCode.
- GLM 5.2: Scored 94 on both, but Hermes consumed significantly fewer tokens than OpenCode.
- DeepSeek V4 Pro: Achieved the highest score of 97 with Hermes, while OpenCode was faster and extremely cheap ($0.06).
The test reveals that pairing models with the right agent framework drastically impacts speed, token efficiency, and accuracy.
More from coding & agent
- MCP Protocol Breaking Update: A Guide to Stateless Migration — lee_stott · 2026-07-31
- LangChain Founder Recommends Ecosystem Stack for Building Agent Harnesses — hwchase17 · 2026-07-31
- Chip Huyen Demos Workflow Orchestrating 1,409 Agents — hugobowne · 2026-07-31
- Agent Gave Boss Fake Link: Dev Reflects on RAG Best Practices — gogeta7124 · 2026-07-31
- DeepSeek-V4-Flash Agent Eval: Completes 3D Task for $0.07 — cedric_chee · 2026-07-31
- Chamath Predicts Agentic Code Will Reduce Software Error Rates and Security Incidents — iamKierraD · 2026-07-31