How to run agent evals against real dependencies without issuing 40 refunds
mangoavococo · reddit · 2026-08-20
A developer asks on Reddit how to set up evals when they must run against dependencies from real app scenarios—feature flags, real traffic, multiple services. When testing agents that call multiple real tools and APIs, you can't have an eval actually issue 40 refunds or print 60 return labels. The thread seeks engineering practices for balancing production dependencies against safe sandboxing.
More from coding & agent
- Grok lets you create agent skills via screen recording — KyeGomezB · 2026-08-20
- terminal-code brings VS Code inside the terminal, works over SSH — rakyll · 2026-08-20
- Built a Local Mac App Where Claude Acts Automatically on Events — KlassyCoder · 2026-08-20
- Cursor Launches Origin, a GitHub Rival for Code Hosting with Agent-Native Features — emmanuelvivier · 2026-08-20
- Client wanted local Claude for all staff; consultant suggested cloud agents — verrsane · 2026-08-20
- Claude + Runway via MCP: One Prompt Writes, Designs and Shoots a 60-Second Film — aziz4ai · 2026-08-20