Generative Agents: Interactive Simulacra of Human Behavior
Joon Sung Park, Joseph C. O'Brien, Carrie J. Cai, Meredith Ringel Morris, Percy Liang, Michael S. Bernstein
cs.HC, cs.AI, cs.LG
2023-04-07
Memory, reflection and planning on gpt-3.5 let 25 Smallville agents spread a party invite; full architecture TrueSkill 29.89 vs 21.21 with no memory.
A large language model can already play a character for one scene. The hard part is keeping a whole town coherent for days: who invited whom yesterday, not eating lunch twice, deciding whether to stop and talk. Finite-state machines and behavior trees cannot cover an open world. Reinforcement learning needs a reward for "acting like a person," which nobody has. Park and colleagues at Stanford HCI, with Percy Liang and Google Research / DeepMind, built an interactive town to test whether an architecture around an LLM can carry that load.
Everything turns on a memory stream, a timestamped log of natural-language observations. At decision time the system does not dump the whole history into the prompt. It ranks memories on three scores:
The three scores are min-max scaled and summed with equal weights; the top memories that fit the context window go into the prompt.
Raw observations are not enough. Asked who to spend an hour with, Klaus would pick Wolfgang, the dorm neighbor he passes in the hall, instead of Maria, who shares his research obsession. Reflection fires when the importance scores of recent events sum past 150, about two or three times a day. The model proposes high-level questions from the last 100 records, retrieves, and writes more abstract insights with citations. Reflections can cite other reflections, so the tree grows from events up to a self-concept.
Planning keeps the day from looping. Ask "what now" and Klaus eats lunch at 12:00, again at 12:30, again at 1:00. The agent first sketches five to eight blocks for the day, then recursively splits them into hour-long and 5-to-15-minute actions. At each step it decides whether to react and replan. Dialogue is conditioned on each side's summary of the other. The backbone is ChatGPT gpt-3.5-turbo; GPT-4's API was still closed.
Smallville is a Phaser sandbox with 25 agents, each seeded by a semicolon-delimited paragraph. Users can rewrite object state (the stove is burning) or speak as an agent's inner voice.
One hundred Prolific raters watched a replay of an agent's two-day life, then ranked five interview answers for believability. TrueSkill converted those ranks into scores:
| Condition | TrueSkill μ |
| Full architecture | 29.89 |
| No reflection | 26.88 |
| No reflection or planning | 25.64 |
| Crowdworker role-play | 22.95 |
| No observation, planning, or reflection | 21.21 |
Full architecture versus the no-memory baseline: Cohen's d = 8.16. Kruskal-Wallis H(4)=150.29, p<0.001. Every pairwise gap is significant except the two last-place conditions.
Over two simulated days, news of Sam's mayoral run moved from 1 agent (4%) to 8 (32%). Isabella's Valentine's party moved from 1 to 13 (52%). Every "yes, I know" traced back to a real memory. Mutual-acquaintance network density rose from 0.167 to 0.74; 6 of 453 "do you know X" answers (1.3%) were hallucinations. Five of twelve invitees showed up; three had conflicts, four said they wanted to come but never put it on the day's plan.
This is the template a long line of social-simulation agents still copy. For game NPCs, social-product prototypes, or rehearsal of hard conversations, you can seed one intent and let diffusion and coordination grow, instead of scripting 25 people. The bill is concrete: two days for 25 agents cost thousands of dollars in tokens and wall-clock days. That is a 2023 gpt-3.5 invoice. The split of labor among retrieval, reflection, and planning is the part that stuck.
The full architecture also beat crowdworkers who had watched the same replay. "Looks like a person" already cleared that floor. The paper does not claim it cleared expert performance.
Retrieval misses: Rajiv had heard about the election and still said he was not following it. Tom remembered he was supposed to discuss the election at the party, and was unsure the party existed. Hallucinations are mostly embroidery, not invented biographies, but Isabella added a "he's announcing tomorrow" that Sam never scheduled. World knowledge leaks: a neighbor named Adam Smith gets credited with Wealth of Nations.
Physical norms do not travel well in language. Agents walk into a one-person dorm bathroom that is already occupied, and into shops after the 5 pm closing. Instruction tuning makes them overly polite; Isabella almost never refuses party ideas that do not fit her, and other people's interests start to look like her own. The evaluation covers two game days. Crowdworkers are not an expert ceiling. Robustness is untested: prompt injection and planting a fake memory through chat are open questions. Biases in the base model pass through unchanged.