Why User-Model Evals Are Hard: Stanford Researchers Bet on a Multi-User Turing Test

alexisjross · x · 2026-10-06

Researcher alexisjross uses the ELIZA case study to highlight two challenges in evaluating user models: (1) even humans are poor discriminators of whether a simulated-user model is convincing, and (2) existing evals require predefining evaluation dimensions, which is unreliable when humans themselves can't reason well about what makes humans human.

Her team instead settled on a multi-user Turing test (MTT) using an LM judge with no priors about good user models — the judge infers the most relevant behavioral features from in-context examples of real vs. simulated users. Noah Goodman notes he's teaching ELIZA this week in Stanford's Minds and Machines course.

Original post →

More from Research

Research channel →