Keeping LLM agents alive on multi-day jobs: heartbeat boards, epoch fencing, and open problems

milkygirl21 · reddit · 2026-10-10

The author runs a fleet of LLM agents on jobs lasting hours to days (e.g. batch transcribing and summarizing hundreds of videos into a knowledge base). The core pain: agents die quietly without crashing. Nine practices already in place:

Open problems: credential delivery (vault only fills browser inputs, so git commits go through a slow browser route), workspace resets wiping local state, no external dead-man switch if the audit and board die together, step/context limits causing dirty stops, and Sheets quota throttling. The author asks for suggestions on all of these.

Related event: Keeping Long-Running Agents Alive: Nine Practices and Five Open Problems(2 posts)→

Original post →

More from coding & agent

coding & agent channel →