What an agent must save so a failed run can be safely replayed and resumed
tomibrumen · reddit · 2026-09-28
The author distinguishes two scenarios when implementing agent workflow recovery: reproducing a failure requires preserving inputs, prompt and tool versions, tool responses, and approval decisions; resuming after a fix means recovering from a safe checkpoint and checking which external actions have already occurred.\n\nThe hard part is that an automatic retry must not duplicate side effects (double-sending messages, creating duplicate records), which requires tracking external IDs and completed actions. The author also balances recovery against retention: keeping enough context to understand decisions without storing every sensitive payload. Three operational responses to dependency failures are proposed: auto-retry from the checkpoint, notify a human with the failure trace, or wait for a human fix and approved rerun.
More from coding & agent
- Swarms creator Kye Gomez: multi-agent systems are AI's most impactful frontier — KyeGomezB · 2026-09-28
- Deel Launches Akai, an Enterprise AI Agent Platform, After Hitting $140M ARR — testingcatalog · 2026-09-28
- AI Agents Now Move Real Capital in Finance, but Audit Trails Haven't Kept Up — Master-Sprinkles-848 · 2026-09-28
- H Company releases Holo4 open VLMs for computer-use agents — jacek2023 · 2026-09-28
- Our RAG bot answered "accept all cookies" as competitor pricing: scraping tool shootout — Typical-Code-7006 · 2026-09-28
- Two prompts: Claude Opus 5.5 codes, renders and self-checks a full 1080p promo video — sujingshen · 2026-09-28