Replay forks make SocioVerse 2 counterfactuals attributable; best ABM score is 0.900 vs 0.911

SocioVerse2: A Longitudinal Dynamic Social Simulation Framework under a Human-AI Co-evolutionary Paradigm

Xinnong Zhang, Jiayu Lin, Jia Wang, Yixu Huang, Xinyi Mou, Yingqian Wu, Jingcong Liang, Shijun Lei, Jianing Shi, Guanying Li, Siyuan Wang, Hanjia Lyu, Zhenfei Yin, Yunlu Yin, Siming Chen, Yulan He, Jiebo Luo, Xuanjing Huang, Liyin Jin, Baohua Zhou, Hanqi Yan, Zhongyu Wei

cs.CL, cs.CY

2026-09-22

SocioVerse 2 replays history up to an intervention, then forks. On ten ABMs, DeepSeek-V3 scores 0.900 consistency versus a 0.911 rule-rerun ceiling.

What problem this solves

A simulation works as an experiment only when a condition can change and a control can stay. Sandboxes such as Generative Agents, Concordia, and OASIS reproduce a known pattern, then freeze the trajectory. Alignment platforms such as HiSim, AgentSociety, and SocioVerse 1.0 match a population to a real society in cross-section, so a new population means a new data collection. YuLan-OneSim and AgentSociety 2 hand the workflow to an autonomous agent, and the researcher mostly approves whatever that agent chooses to show. Audits of this pattern find made-up results and drift from the stated plan.

SocioVerse 2, from Fudan and collaborators, turns version 1.0 into two loops. One loop moves a population through time and forks counterfactuals. The other treats the study design as state that can be edited. Population and environment sit in separate services.

Method

Behavior is B = f(P, E). Population, environment, and the behavior function can be swapped on their own. The function may be a classical rule, an LLM prompt, a reinforcement-learning policy, or a mix. Collective behavior rewrites the next environment. The study state also holds metrics, a resource manifest, and a seed. An edit on any piece produces the next version.

Each step has five phases. Scheduled events and broadcasts fire first, and a declared intervention enters in that same phase. Every agent then sees a slice of the world, split into physical versus information and broadcast versus direct message, and returns a typed action with a source label. At the end of the round the engine writes the panel. Memory defaults to eight steps and can be rebuilt from that panel, which is why a rule and an LLM can share one action space.

Counterfactuals are replayed. An intervention δ = (t, op) edits the environment at step t. Earlier actions are copied from the parent panel, and the decision model is not called. From t onward the model decides again. Population, behavior function, and seed stay fixed, so the later gap belongs to that single edit. A branch changes only the environment and keeps history. A version may change population, behavior, or data, keeps no history, and snapshots the workspace so the parent can be restored.

The population MCP, a typed interface, indexes 10,448,375 personas locally: 973,928 from X and 9,158,404 from Xiaohongshu. Iterative proportional fitting reweights the draw onto census marginals. Cells no pool covers are synthesized under a synthetic label, or dropped with the gap reported. The event MCP serves more than 21 signal sources and, for a given year and month, returns only data already public then. A run reads the bundle stored in the workspace. Skills pause for a person by default. Auto mode can run straight through.

Results

The mechanism cases are the ones with a full numeric control. Ten canonical agent-based models, 40 to 810 agents and 30 to 200 steps, cover flow, markets, organizations, and diffusion. An agent-based model lets local rules produce a group pattern. Consistency is one minus the root-mean-square relative deviation from the rule reference. Ten reruns of the rule average 0.911, the noise ceiling under this protocol. Three LLMs use three seeds each, temperature 0.7, at most 256 tokens.

MethodMetricResult
Rule controlMean consistency, 10 tasks0.911
DeepSeek-V3Mean consistency, 10 tasks0.900
GPT-4oMean consistency, 10 tasks0.898
Qwen3-235BMean consistency, 10 tasks0.885

DeepSeek-V3 at 0.900 sits 0.011 under that ceiling. The three models lie within 0.015 of one another. Task structure moves the score more than model choice.

On Schelling, a 20 by 20 grid with 400 agents and threshold 3/8, scalar neighborhood statistics send the LLM to full satisfaction at step 12, on the rule's path. Natural language plus moving cost and patience keeps the fully segregated end state and delays it to step 18. Context changes the tempo.

Prisoner's Dilemma is the miss. The rule control is nearly self-identical, all three LLMs land well below it, and the prose prints no score for that task. Continuous headings record every small error. On opinion dynamics the rule itself posts the suite's lowest score, and the LLMs are not worse there. On the minority game, GPT-4o and DeepSeek-V3 beat the rule control.

The media study uses real participants. Each movement has 1,000 people, 300 LLM core users and 700 rule-based users, across 14 steps of 12 hours. A mirror agent copies the core stance into the rule layer. Influence does not flow back. Stance accuracy is 0.968 on MeToo, 0.943 on Roe v. Wade, and 0.899 on Black Lives Matter, with stance F1 at 0.340, 0.336, and 0.374. Content similarity runs from 0.806 to 0.841, content-class accuracy from 0.642 to 0.735. Behavior accuracy runs from 0.667 to 0.780, behavior F1 from 0.469 to 0.576.

Five further cases are each a single run. Chicago segregation uses 19,235 households over 15 steps against the 2010 Census and five rule models. Drug procurement uses 325 markets over 50 rounds against historical winners, an enterprise-only policy, and random. Consumer confidence uses 250,000 households over 75 months against the official index and 12 baselines. The purchasing managers' index uses 300 firms over 22 plus 32 months against the official series, persistence, and consensus. German car shares use 600 consumers over 41 months against registrations.

Why it matters

The piece to reuse is the version tree. A swap of population, behavior function, or data vintage can be rolled back to the parent. Steps before a fork are panel replay, with no further model calls. By design the LLM drives a few hundred core users.

The 0.011 gap is a reproduction. Where actions are discrete and a rule already explains the mechanism, an LLM lands inside the rule's own rerun noise. The abstract also places macro nowcasting past the response model's knowledge cutoff: estimate the current official index from signals already public. The ABM gap does not carry that claim.

Limitations

The suite mean hides the tasks that break. Prisoner's Dilemma and continuous control have no per-task scores in the prose, and the LLMs use 3 seeds against the rule's 10. A gap of 0.011 is not a stable margin.

Stance accuracy and F1 disagree. An accuracy of 0.968 beside an F1 of 0.340 is what a majority class does. Micro alignment reported through accuracy alone is overstated. The mirror link is one-way, carried over from the earlier design. Five applied cases run once, so a monthly path that tracks an official series mixes mechanism with a single draw. Human edits inside the two loops are not compared with a fully autonomous run of the same specification, and auto mode can skip the checkpoints. Xiaohongshu accounts for most of the locally indexed personas. MatrAIx is calibrated on one-dimensional marginals only and represents no real population. Synthetic cells, if later read as real respondents, inflate external validity.

Terms

Source

What people are saying

Related papers

All paper explainers