Gemini 3.7 Flash Tops Zapier's AutomationBench, Beating Pricier Claude and GPT Models
_philschmid · x · 2026-08-14
Zapier's AutomationBench evaluates LLMs on end-to-end workflow execution using 47 real tools across six business functions: Sales, Marketing, Operations, Support, Finance, and HR.
Gemini 3.7 Flash ranks 1st with a 30.44% success rate and a cost of $0.61 per task, outperforming more expensive models like Claude Opus 5 and GPT-5.6 Terra. It dominates in Marketing, Finance, Sales, and Support, but trails behind Claude Opus 5 in Operations.
More from coding & agent
- Databricks Introduces Smart Routing in Unity AI Gateway, Claims 30%+ Cost Reduction — matei_zaharia · 2026-08-14
- DAB Benchmark: Simulating Messy Data Warehouses to Expose AI Agent Flaws — HamelHusain · 2026-08-14
- Developer Reports DSH Framework Shows High Efficiency in Complex Projects — ChrisGPT · 2026-08-14
- Entire drops waitlist: Git hosting for AI coding agents now open to all — craigsdennis · 2026-08-14
- Prime Agent Open-Sourced: A Recursive Agent That Rewrites Its Own Prompts — JeremyCMorgan · 2026-08-14
- Toast 1 Search Agent Tested: Frontier Quality at 1/10th the Price — xeophon · 2026-08-14