How to run agent evals against real dependencies without issuing 40 refunds

mangoavococo · reddit · 2026-08-20

A developer asks on Reddit how to set up evals when they must run against dependencies from real app scenarios—feature flags, real traffic, multiple services. When testing agents that call multiple real tools and APIs, you can't have an eval actually issue 40 refunds or print 60 return labels. The thread seeks engineering practices for balancing production dependencies against safe sandboxing.

Original post →

More from coding & agent

coding & agent channel →