G2 launches agent evals: 10 enterprise support agents score 56%-92% on policy compliance across 46 scenarios
sanderssays · x · 2026-09-15
G2 launched agent evals: 10 enterprise support agents run against the same simulated company across 46 scenarios, with policy compliance ranging from 56% to 92%. The author previously co-founded Mosaiaio building agents for investment teams, where clients needed weeks of demos to know what to automate — so they built independent benchmarking: each agent is configured like a real CX team would, then runs the same support desk. Sales and legal categories are next, expanding to 10 total.
More from coding & agent
- Claude Mods: Anthropic to ship function hooks-based plugins for Claude Code in weeks — bcherny · 2026-09-15
- The hidden cost of failed agent runs: one task may burn twice the credits on Codex — ToneAromatic178 · 2026-09-15
- OpenAI Agents API enters public beta: managed cloud agents on the Codex harness — craigsdennis · 2026-09-15
- Interview with the Claude Code team on building it while models keep outpacing engineering — EricBuess · 2026-09-15
- Before buying a faster model, check your agent's traces: serial API calls eat latency — gethackteam · 2026-09-15
- ElevenLabs MCP adds voice, music, image, and video generation — lukeharries · 2026-09-15