DeepSeek V4 Flash beat GLM 5.2 and Kimi K3 on a multi-app agent benchmark
LimpComedian1317 · reddit · 2026-08-04
DeepSeek V4 Flash beat GLM 5.2 and Kimi K3 on a hard agentic benchmark
Composio says it tested DeepSeek V4 Flash, GLM 5.2, and Kimi K3 on difficult long-running agent tasks that span apps like PagerDuty, Gmail, HubSpot, Airtable, and Slack.
- The harness used Pi Agent + Composio MCP to keep the evaluation impartial.
- DeepSeek V4 Flash was the fastest, at 164 seconds per task — about 2.5× faster than GLM and 1.4× faster than Kimi.
- Average task cost was roughly $0.08 for DeepSeek, versus $0.57 for GLM and $1.39 for Kimi.
- Success rates were close: 20/30 for DeepSeek, 21/30 for both Kimi and GLM.
- Frontier models still led the suite: Fable 5 and GPT-5.6 Sol scored 24/30, while Opus 5 scored 23/30.
- The models also showed different working styles: GLM was token-efficient, DeepSeek was fast but token-hungry, and Kimi sat in between.
The post asks for others’ experience with open-weight models in agentic workflows.
More from coding & agent
- GitHub Stacked PRs repo shows how to split one big review into layered changes — DanWahlin · 2026-08-04
- Evedev is being recommended as the default framework for internal agents — cramforce · 2026-08-04
- Agent swarms fail when they optimize for process instead of shipping code — doodlestein · 2026-08-04
- A month-by-month meme tracks how coding-agent habits keep changing — unixterminal · 2026-08-04
- A critique says OpenCode’s agent design breaks KV cache and weakens security — JFPuget · 2026-08-04
- YC-backed Buildbox launches agent analytics for real user outcomes, not just evals — ycombinator · 2026-08-04