2026-10-10
Twins built from 500+ answers of 1,784 people scored 0.748 accuracy on 164 outcomes, only 0.014 above an empty prompt, with mean correlation r=0.20.
People are already using large language models as stand-ins for specific individuals: filling surveys, running experiments, rehearsing policy. The early numbers look strong. Twins built from two-hour interviews reached about 85% accuracy relative to test-retest on the General Social Survey. The public Twin-2K-500 panel reported 72% accuracy on held-out items, 88% relative to test-retest.
Neither number travels well. The interview data are not public, and the distance between the interview and the test items is unclear. Twin-2K-500 was not compared with demographics-only personas or with an empty prompt, and the twins reproduced only about half of the human experimental effects. Many validation items were classic published paradigms, which the base model may have seen.
The core group is at Columbia, with 23 coauthors each bringing a question from their own work. The base resource is Twin-2K-500: more than 2,000 U.S. participants, each with answers to more than 500 questions, about 128,000 characters. The mix is 14 demographic items, 279 items from 19 personality batteries, 85 cognitive items, 34 economic-preference items, 48 items from heuristics-and-biases experiments, and 40 pricing items, collected in four waves over about 145 minutes on average. Those human answers retest reasonably well and largely reproduce known effects.
A twin is built by in-context learning. The person's question-answer record goes into the prompt, and the model answers a new survey. The default base model is GPT-4.1 dated 2025-04-14, temperature 0.7. Randomization in the survey is preserved so the twin sees the version that human saw. Outputs have to land inside the allowed range; failures are retried.
Nineteen preregistered sub-studies, 164 outcomes. Cumulative responses 13,506, unique participants 1,784. The topics include creativity, public-issue attitudes, privacy, fairness, luxury goods, news, and labor-market preferences, mixing published paradigms, unpublished ones, and new designs, often with novel stimuli. Humans answered on Prolific first. Each twin then answered the same items, matched person to person.
Two preregistered metrics carry the paper. Individual-level accuracy is 1 minus the mean absolute deviation divided by the outcome's range, so 1 is a perfect match. Open-ended items without a fixed range are dropped, leaving 161 outcomes. Correlation is computed across participants within each outcome, averaged on the Fisher z scale, then transformed back. They also report the gap in means as Glass's delta, and the ratio of standard deviations.
Four benchmarks sit underneath. Uniform random answers on the outcome's range. An empty persona, the same prompt for everyone with no personal information. A demographics-only persona using the 14 demographic variables. And, later, XGBoost trained on the same features. They also varied temperature, and swapped in GPT-5, Deepseek, Gemini including Gemini-3 Pro, a GPT-4.1 fine-tuned on Twin-2K-500, and the behavior model Centaur.
Full-persona twins average 0.748 individual-level accuracy. Random answers already score 0.629, the empty persona 0.734, demographics only 0.746. The full persona beats the empty persona by 0.014, and that gap is significant. The 0.002 edge over demographics is not, p=0.37. A score of 0.748 looks respectable because, on a bounded item, guessing near the middle is far from zero.
Mean correlation across people is 0.197. The correlation is positive for 157 of 164 outcomes, and significant after Bonferroni correction for 97. Demographics-only personas reach 0.145, empty personas 0.080, random answers 0.001. An r of 0.20 is about the correlation between height and intelligence. Earlier interview-twin papers reported higher correlations by correlating across questions within a person. This paper correlates across people within a question. That is the comparison a product or a study usually wants: who is more likely to buy, who is more likely to agree.
On means, twins differ from humans by 0.352 standard deviations. The paper's own conversion puts the error of that mean estimate near the error of a sample mean from about 10 humans. The difference is significant for 105 of 164 outcomes, 64%. Twin standard deviations are lower than human ones in 154 outcomes, 93.9%, and 140 of those gaps are significant.
| Method | Individual accuracy | Correlation with the person |
| Full persona, temperature 0.7 | 0.748 | 0.197 |
| 14 demographics only | 0.746 | 0.145 |
| Empty persona | 0.734 | 0.080 |
| Uniform random | 0.629 | 0.001 |
| GPT-4.1, temperature 0 | 0.752 | 0.232 |
| Full-persona XGBoost, 650 training people | No single summary given | Still below 0.29 |
Across temperatures and base models, the best setting is GPT-4.1 at temperature 0: correlation 0.232, accuracy 0.752. Compressing the 128,000-character record into a statement-style summary of about 13,000 characters performs about as well as the full persona. Centaur as the base model underperforms GPT. The paper leaves the reason open: the Llama base, the formatting, or an out-of-distribution test.
Five distortions show up in the same matched data.
Insufficient individuation. Mean absolute deviation between full-persona twins and empty-persona twins is 0.175. Between full-persona twins and the humans it is 0.252. The extra record moves the answer off the base model, and the move is still smaller than the remaining gap to the person.
Stereotyping. Mean absolute deviation between full personas and demographics-only personas is 0.132, smaller than the gap to the empty persona and the gap to humans. Correlation does improve. On one "lack of control" item in sub-study 2, demographics-only correlation is 0.105 and full-persona correlation is 0.555, while accuracy moves only from 0.892 to 0.907. The twin gets better at sorting people and stays about as far from the score that person actually gave.
Uneven representation. An XGBoost model predicts each participant's accuracy from 61 demographic dummies. Partial dependence plots show higher accuracy for people with more education and higher income, and for people with moderate positions on the party spectrum and middling attendance at worship.
Directional bias in the content. Besides the 164 detailed outcomes, 31 higher-level comparisons were preregistered. Relative to their humans, twins lean toward a rosy view of people: more likely to say others are fair and trustworthy, more favorable to self-reliance, more favorable to donors who give to both parties, more supportive of regulating add-on fees, and more willing to pay taxes for everyone's healthcare. They also lean toward technology: more accepting of algorithmic hiring, less bothered by online targeting, and they under-report Netflix and TikTok use. When rating creativity, twins score ideas lower when the source is a human.
Hyper-rationality. On items with an objective answer, twins are close to omniscient. On novel paradigms they fail to show attraction effects, compromise effects, and default effects. On a classic published default-effect paradigm, they show a strong default effect. The paper reads the second pattern as leakage: the paradigm is famous, and the answer may already sit in the base model's training text.
Sixty-eight experts, mostly researchers and practitioners, guessed mean treatment outcomes in five sub-studies. The human average treatment effect fell inside the experts' 95% interval all five times. The twin average treatment effect fell outside it three times.
Against a model that is allowed to see real answers, the comparison uses the 106 outcomes with sample size above 650. XGBoost is trained per outcome on the same full-persona features plus that outcome's answers for a training subset, from 50 people up to 650. At 650, correlation is still below 0.29. Full-persona twins, which never see the outcome, match roughly an XGBoost model trained on about 180 people's answers for correlation, or about 225 in the demographics-only case. On accuracy, twins match roughly an XGBoost model trained on about 75 labeled answers.
By domain, correlations are higher for cognitive items, human-technology interactions, and scale questions, and also for conflict, prosocial topics, social cognition, and personality. They are lower when social desirability is salient, on elections and public issues, on valenced judgments, and when the question varies across participants.
Anyone about to replace a survey with synthetic respondents should read 0.748 against its baselines. It is 0.014 above an empty prompt. Sorting people works a little, at r about 0.20. Five hundred questions do not pin down what this person will say.
The ceiling is the sharper result. Give the same record to XGBoost and also show it up to 650 real answers on the item, and correlation still stays under 0.29. The shortfall is not only that a language model fails to use the features. The 500 questions, drawn from decades of standard scales, do not predict answers in a new setting. GPT-5 and Gemini-3 Pro, in the settings they tried, do not change that picture.
The usable case is narrow. A twin is closer to a comparative profile colored by the base model. It can roughly rank who is more open to some arrangement. It is a poor clone of the person. Treating it as a well-informed, unusually rule-following advisor is more honest than treating it as a copy. The data and code are public. The Hugging Face release, Twin-2K-500-Mega-Study, is a testbed for the next pipeline.
Only one construction was tested: a public questionnaire in the context window. Interview twins and the fine-tuning pipelines inside closed products were not in the comparison. The defense is that Twin-2K-500 has about 15,000 downloads and is the most replicable public setup. The result lands on that setup, not on every product that uses the name digital twin.
The five distortions are exploratory. The metrics and the 31 higher-level comparisons were preregistered. The claim that the failures take exactly these five forms was not. The directions match earlier synthetic-data work. The main text illustrates the content biases with selected sub-studies, and does not print an effect size for all 31 comparisons. "Pro-person and pro-technology" is a pattern they flag, not a finished ledger.
Individual-level accuracy is easy to quote as "already 75%." The random baseline is 0.629 and the empty model is 0.734. The increment is the result. The number to cite is the correlation, 0.20.
The sample is the U.S. Prolific panel behind Twin-2K-500. Outside groups cannot recontact the same people, because platform IDs are not shared. They can only merge new analyses onto the released answers. Centaur's weaker score is not, by itself, evidence that a behavior-tuned base model is useless. Format and distribution shift were not isolated.