Top agents drop from 58% to 35% on multi-turn CRM tasks, with near-zero confidentiality awareness
Shahules786 · x · 2026-09-03
Thread (3/6): CRMArena-Pro results show why the setting matters — leading agents reach 58% success on single-turn tasks but fall to 35% when users reveal information over multiple turns. Models show near-zero inherent confidentiality awareness; prompting improves refusals but often hurts task performance. The hard part is gathering missing context, applying policy, and knowing when not to answer.
More from coding & agent
- Fable 5.1 produced a result in ~10 minutes on the $100 plan — jasondeanlee · 2026-09-03
- Hermes Agent v0.21.0 Adds Persistent Multi-Gateway Connections for Desktop — Teknium · 2026-09-03
- 7 GitHub Repos Turn One AI Agent Into an OS: Memory, Model Routing, Free Compute Stacks — garrytan · 2026-09-03
- The Zvi: agents rewire your reflexes — annoyances now get fixed by just asking Claude Code — TheZvi · 2026-09-03
- "A smarter model in a bad system just makes expensive mistakes faster" — iamKierraD · 2026-09-03
- Omnara: open-source, self-hostable alternative to Claude managed agents — JaynitMakwana · 2026-09-03