Salesforce Test: AI Agents Succeed on Only 58% of Real CRM Tasks
aakashgupta · x · 2026-10-05
Citing 2025 Salesforce research, the author notes that leading AI agents doing realistic CRM work succeeded about 58% of the time on single-turn tasks, dropping to roughly 35% once the task became multi-turn. The post argues benchmarks are tidier than real companies: the same customer appears twice in the CRM, Slack escalations contradict tickets, and APIs behave inconsistently — the messiness that actually breaks agents. Takeaway: real-world enterprise reliability lags far behind leaderboard scores.
More from coding & agent
- marimo Studio: one notebook, separate audience views, agent-safe presentation layer — pandeyparul · 2026-10-06
- AI software factories: developers state intent, agents handle build and deploy — Pavan_Belagatti · 2026-10-06
- Dev builds polished card animations through multiple rounds with Astra — Dimillian · 2026-10-06
- pg-jev open-source Postgres extension runs semantic filters and scoring inside SQL — Arindam_1729 · 2026-10-06
- Open-source Discord AI assistant Zauq v4 ships bounded agents, MCP and sandboxed code verification — rar_file-exe · 2026-10-06
- Orbio launches new agent tools: sandboxes, 24/7 servers, databases and email, paid in CREDIT — econoar · 2026-10-06