Testing shows GLM agent teams turn easy tasks into slop, but the model shines at swarm science

teortaxesTex · x · 2026-09-11

Blogger teortaxesTex warns against using DSH agent teams on projects without clear modularity: Astra's audit was brutal, and V4.1 subagents turned a task that wasn't hard for V4.1 alone into slop. Yet the quoted tweet also shows strengths: asked to build a nacre fan, V4.1 thought critically about material properties, downloaded 180MB of PDFs and assigned a subagent to review the physical optics of pearls — the author deems it already fit for swarm scientific work, while GLM 5.3 Flash just thinks about three.js.

Related event: Hands-on warning: DSH agent teams underperform without clear modularization(2 posts)→

Original post →

More from coding & agent

coding & agent channel →