Running an LLM town with 800+ persistent agents: concurrency, caching, and costs
Low_Bad_6585 · reddit · 2026-10-08
An indie developer's engineering write-up on Slow Vale, an LLM life simulation whose Chinese server runs 800+ persistent AI residents in one continuously running city. Key points: each character makes 300–400 LLM calls/day at 30k average context; inference runs concurrently but responses go through validity checks before mutating world state (the last fish on a shelf may be gone by return time); a shared runtime manages activities with different durations, resource conflicts, and recovery; decision contexts mix information with different update frequencies (stable personality vs. second-level inventory changes); hosted DeepSeek is used rather than local inference.
More from coding & agent
- Hack: run coding agents in Grok's cloud computer to burn unused Code Plan credits — op7418 · 2026-10-08
- Grok bot surfaces $25,730 in live GitHub bounties, top one pays $10k — prasenx · 2026-10-08
- Talk: how to RL-train an agent running inside a harness you didn't write — SergioPaniego · 2026-10-08
- 8 open-source AI agent tools to know, from computer-use to browser automation — nikola_mr64990 · 2026-10-08
- Test quality follows module design: test larger units instead of banning AI tests — mattpocockuk · 2026-10-08
- Matt Pocock: AI test quality depends on codebase design, banning tests is wrong — mattpocockuk · 2026-10-08