Agent finished the task, but the caller never got the answer: a restart incident postmortem

jonah_omninode · reddit · 2026-10-02

OmniNode shares a real incident: a task arrived one second before a redeploy restarted services. The agent finished its work 23 seconds later, but the caller timed out after five minutes with no answer.

Two root causes: the request was marked handled before returning, and the shutdown sequence closed the reply connection before flushing queued answers, leaving the caller unable to tell failure from in-progress.

After the fix, graceful restarts return an explicit failure, and hard restarts redeliver the request to completion. But the author warns that redelivery can repeat side effects like file edits or comments, since idempotency checks aren't built yet.

He asks peers how callers should learn a task's outcome when workers restart mid-flight.

Original post →

More from coding & agent

coding & agent channel →