Running an LLM town with 800+ persistent agents: concurrency, caching, and costs

Low_Bad_6585 · reddit · 2026-10-08

An indie developer's engineering write-up on Slow Vale, an LLM life simulation whose Chinese server runs 800+ persistent AI residents in one continuously running city. Key points: each character makes 300–400 LLM calls/day at 30k average context; inference runs concurrently but responses go through validity checks before mutating world state (the last fish on a shelf may be gone by return time); a shared runtime manages activities with different durations, resource conflicts, and recovery; decision contexts mix information with different update frequencies (stable personality vs. second-level inventory changes); hosted DeepSeek is used rather than local inference.

Original post →

More from coding & agent

coding & agent channel →