What an agent must save so a failed run can be safely replayed and resumed

tomibrumen · reddit · 2026-09-28

The author distinguishes two scenarios when implementing agent workflow recovery: reproducing a failure requires preserving inputs, prompt and tool versions, tool responses, and approval decisions; resuming after a fix means recovering from a safe checkpoint and checking which external actions have already occurred.\n\nThe hard part is that an automatic retry must not duplicate side effects (double-sending messages, creating duplicate records), which requires tracking external IDs and completed actions. The author also balances recovery against retention: keeping enough context to understand decisions without storing every sensitive payload. Three operational responses to dependency failures are proposed: auto-retry from the checkpoint, notify a human with the failure trace, or wait for a human fix and approved rerun.

Original post →

More from coding & agent

coding & agent channel →