Ben Hylak on building simulations for agent evals: replay traces, detect sim awareness
HamelHusain · x · 2026-10-11
Hamel Husain interviewed ben hylak about implementing simulations for agent evals across several companies.
Key points:
- Why input/output tests miss agent failure modes
- Replaying production traces to see what a change breaks
- Recreating the data and tools around the agent
- 'Simulation awareness': models detect they're in a simulation and change behavior
- Choosing scenarios that actually test your change
- Checking answers, costs, and tool behavior across runs
- Replaying known failures to verify fixes
- Common mistakes that make simulations misleading, and validating that sims reproduce production
Related event: Hamel Husain Interviews benhylak on Simulation-Based Agent Evals(3 posts)→
More from coding & agent
- Agentic evals: a practical guide to knowing whether your AI agent did the job — and will do it again — blaizedsouza · 2026-10-11
- Ten 'make a DSL' prompts: a dev's trick for turning AI into a language designer — round · 2026-10-11
- Dev builds personalized fitness app with Grok Code in one day — kevinkern · 2026-10-11
- Heavy User Publishes Claude Code Field Guide: Prompts Down Two-Thirds, Output Up Eightfold — AaronBergman18 · 2026-10-11
- Gaussian splatting + webcam makes websites look 'too real', raising privacy questions — RileyRalmuto · 2026-10-11
- Claude Opus 5.5 decompiles so well that Halo, GTA and CoD now run in browsers — SuB8u · 2026-10-11