Amazon AGI Lab: Why 80% Reliable Agents Still Break in Production
AI Engineer · youtube · 2026-10-09
Felipe Blanes of Amazon AGI Lab shared lessons from serving customers of Nova Act, Amazon's browser-agent service, at AI Engineer World's Fair 2026.
The benchmark illusion: static evals look great until real customers do things nobody expected — then all you have left is hope.
The fix: a customer-driven eval flywheel
- Define success the way the customer does, not the benchmark
- Capture real signals, including talking to customers directly
- Diagnose whether gaps are in the model, harness, or product
- Feed findings into decisions so evals get smarter every week
Key insights
- The trust cliff: at 80% reliability the agent feels like more work; 92% feels trustworthy
- Being open about limits builds more trust than higher benchmark scores
- Surprising use cases from Amazon Leo, Hertz, and Sola customers
- Four principles for keeping evals fresh
More from coding & agent
- Atlassian Launches Agentic Multiplayer Protocol (AMP) for Human-Agent Collaboration — davidhoang · 2026-10-09
- Flock: Open-Source Tool Runs Multiple Claude Accounts Side by Side on Desktop — AccomplishedCraft389 · 2026-10-09
- app-store-screenshots hits 7.2k GitHub stars: AI agent skill scaffolds store-ready screenshot editor — tom_doerr · 2026-10-09
- rhp 2.0 Ships MCP Server So AI Agents Generate Editorial-Style Charts in One Prompt — bezdazen · 2026-10-09
- LLM-as-a-Verifier: Weaker Model Verifies Stronger One, Hits 69.2% SOTA on Terminal-Bench 4 — Azaliamirh · 2026-10-09
- TestSprite Season 4 developer contest offers $4,000 prize pool for Claude Code and Codex users — JaynitMakwana · 2026-10-09