12 real multi-app agent tasks put Fable 5 and Kimi K3 in a tie at 7/12, while GPT-5.6 Sol finished last
Nearby_Pair_6483 · reddit · 2026-07-28
Ran 12 real multi-app agent tasks on live Gmail, Slack, Sheets, Salesforce, HubSpot, GitHub and Linear accounts across Fable 5, Kimi K3 and GPT-5.6 Sol.
- Setup: Claude Code drove Fable and Kimi; Codex CLI drove GPT-5.6, with the same 12 templates and the same MCP tool router.
- Scoring: Fable 5 and Kimi K3 each scored 7/12; GPT-5.6 Sol scored 6/12.
- Cost ceilings at list price: Fable 5 about $7.76/case, GPT-5.6 about $2.69, and Kimi K3 about $1.39; total suite costs were roughly $93, $32 and $17 respectively.
- The hardest tasks were five cross-app reconciliation jobs; all three models failed all five, with one-bad-merge-kills-the-run rules exposing many near misses.
- GPT-5.6 often got closer on partial credit, but the author argues that near misses are dangerous in production.
- For a CRM identity dedup task, Fable and Kimi passed while GPT-5.6 missed 2 checks, which ended up being the decisive gap.
Takeaway: for ordinary SaaS tool use the cost gap may matter more than small score differences, but for exact state reconciliation none of the three should run unsupervised; a verifier plus retry loop is still needed.
More from coding & agent
- Controlled study finds multi-agent setups give 0.0% average gain over a single agent — bravo_abad · 2026-07-28
- Persistent state for coding agents should include approvals, tool traces, and validation results — DesktopLabHQ · 2026-07-28
- Codex Computer Use is showing obvious PMF in compressed computer workdays — hudzah · 2026-07-28
- Splitting planner and executor roles cut my agent context bloat and token bills — truecakesnake · 2026-07-28
- GitHub repo shows how Streamlit theme config helps agent-assisted UI edits — andfanilo · 2026-07-28
- Bilinc launches a hosted MCP memory server to help agents remember across sessions — atakanelik34 · 2026-07-28