New SDK Tests AI Agents on Real Websites and Workflows, Not Mocked Tools
Able-Ingenuity8885 · reddit · 2026-09-25
A team built an SDK for evaluating agents against real websites and workflows instead of mocked tools. Give it a task like 'find a laptop under $1,000, apply the best discount, and check out,' and it pinpoints exactly where the agent succeeds, gets stuck, or silently does the wrong thing. The MVP is done and they are recruiting agent developers for free testing.
More from coding & agent
- The wiki is an anti-pattern: docs must live in the repo for coding agents to maintain — TechPreacher · 2026-09-25
- Microsoft Foundry adds voice experiences for AI agents — lee_stott · 2026-09-25
- Caching policy lookups per inode cuts eBPF security agent CPU cost ~90% — JeremyCMorgan · 2026-09-25
- jev-test-impact uses an LLM to pick which tests to run from a Git diff — alchemist-301 · 2026-09-25
- Astronomer builds his dream learning tool in ~5 hours with Opus 5.5, now live — niloofar_mire · 2026-09-25
- Veteran dev's GSC Wizard MCP does the heavy lifting in ClickHouse, LLM gets conclusions only — DutchSEOnerd · 2026-09-25