12 real multi-app agent tasks put Fable 5 and Kimi K3 in a tie at 7/12, while GPT-5.6 Sol finished last

Nearby_Pair_6483 · reddit · 2026-07-28

Ran 12 real multi-app agent tasks on live Gmail, Slack, Sheets, Salesforce, HubSpot, GitHub and Linear accounts across Fable 5, Kimi K3 and GPT-5.6 Sol.

Takeaway: for ordinary SaaS tool use the cost gap may matter more than small score differences, but for exact state reconciliation none of the three should run unsupervised; a verifier plus retry loop is still needed.

Original post →

More from coding & agent

coding & agent channel →