astra hits 41.4% on automationbench, a new benchmark for agent business-workflow completion
jdjohnson · x · 2026-09-04
Bryan Helmig explains automationbench, where astra set a new high by correctly completing 41.4% of business workflows. The benchmark places models in simulated apps with a job to do, then verifies resulting records and messages with code. Key design challenge: catching subtle errors like updating the wrong customer while letting agents find their own path. Each run starts with a fresh copy of a structured environment (CRM records, emails, spreadsheets, calendar) with API-like tools—nothing touches real customers.
More from coding & agent
- Clanker Cloud Opens Free Web Trial With $20 Worth of Starter Credits — tekbog · 2026-09-04
- Clanker Cloud Lets Anyone Build and Host Agents, With an Enterprise Sales Cautionary Tale — tekbog · 2026-09-04
- Running Grok bots like a company: AI project manager coordinates specialist agents — FinanceYF5 · 2026-09-04
- AI Finds Bugs Faster Than It Fixes Them: Engineers Grapple With CVE Backlogs — _jaydeepkarale · 2026-09-04
- Should Agents Govern Themselves? AAV Adds an External Action-Verifier Layer — CarlosMarreroAAV · 2026-09-04
- Dev spends $40 on classifier evals to cut costs: 'hard to use AI when you can't afford intelligence' — zeeg · 2026-09-04