DeepSeek V4 Flash beat GLM 5.2 and Kimi K3 on a multi-app agent benchmark

LimpComedian1317 · reddit · 2026-08-04

DeepSeek V4 Flash beat GLM 5.2 and Kimi K3 on a hard agentic benchmark

Composio says it tested DeepSeek V4 Flash, GLM 5.2, and Kimi K3 on difficult long-running agent tasks that span apps like PagerDuty, Gmail, HubSpot, Airtable, and Slack.

The post asks for others’ experience with open-weight models in agentic workflows.

Original post →

More from coding & agent

coding & agent channel →