6 models tested on real MCP servers: Opus 5.5 leads, open models cost 87% less per attempt

shensi · x · 2026-10-09

Merge API tested 6 models on the same 30 multi-step tasks using real MCP servers in Claude Code. Frontier models like Opus 5.5 led on completion rate and speed, while open models like MiMo cost 87% less per attempt. The thread examines whether Opus's lead is actually worth the price gap — practical data for agent stack model selection.

Original post →

More from coding & agent

coding & agent channel →