Running a fleet of long-horizon LLM agents: 9 keep-alive patterns and 5 unsolved problems

milkygirl21 · reddit · 2026-10-10

The author runs LLM agents on multi-hour to multi-day jobs (e.g., batch-transcribing hundreds of videos into a searchable knowledge base) and shares a battle-tested keep-alive system: a shared heartbeat board with six standardized fields per job, deadlines at 2-3x check intervals, bounded retries, a parent audit agent sweeping the board every 30 minutes with grouped alerts (mass flags usually mean outage/quota, not dead agents), treating throttled status checks as "unknown" rather than dead, intent/done records for irreversible side effects, epoch fencing for relaunches, drive-first checkpoints, and a watchdog that monitors the auditor itself.

Still unsolved: agents can't safely receive git credentials (vault only fills browser inputs), workspace resets wipe uncheckpointed state, no external dead-man switch if the audit and board die together, step/context limits cause mid-task stalls without handoff, and Google Sheets read quotas get exhausted at this scale.

The author is soliciting concrete suggestions: cheap external dead-man switches, better heartbeat storage, safe secret delivery to agents, clean handoff patterns near step limits, and what to cut.

Related event: Keeping Long-Running Agents Alive: Nine Practices and Five Open Problems(2 posts)→

Original post →

More from coding & agent

coding & agent channel →