CS student asks how shipped agents handle confident-but-wrong actions on real systems

Professional-Mine681 · reddit · 2026-09-29

A final-year CS student scoping a capstone on agent safety recounts his internship project: an agent that scraped funding news and matched leads, where the real risk wasn't dumb outputs but confident wrong actions with access to live systems — junk in the CRM, acting on half-scraped pages, loop-induced API burn. He hand-rolled ad hoc checks each time and found existing guardrail tools (Brane Core, Promptfoo, various SDKs) too action-by-action or prompt-level. He's polling practitioners on worst agent failures, single vs. chained failure modes, current mitigations, and why they abandoned guardrail tools.

Original post →

More from coding & agent

coding & agent channel →