Why User-Model Evals Are Hard: Stanford Researchers Bet on a Multi-User Turing Test
alexisjross · x · 2026-10-06
Researcher alexisjross uses the ELIZA case study to highlight two challenges in evaluating user models: (1) even humans are poor discriminators of whether a simulated-user model is convincing, and (2) existing evals require predefining evaluation dimensions, which is unreliable when humans themselves can't reason well about what makes humans human.
Her team instead settled on a multi-user Turing test (MTT) using an LM judge with no priors about good user models — the judge infers the most relevant behavioral features from in-context examples of real vs. simulated users. Noah Goodman notes he's teaching ELIZA this week in Stanford's Minds and Machines course.
More from Research
- AI eval researcher: even humans can't detect subtle AI patterns like distributional biases — alexisjross · 2026-10-06
- RT-SAFE benchmark: 94.1% of 8 frontier VLMs reach goals, only 0.7% finish with zero safety events — Lianhuiq · 2026-10-06
- Vibe Robotics: Artifact Arena makes frontier models engineer competing robots — kaixhin · 2026-10-06
- Pan-cancer tumor segmentation model wins Best Accuracy at MICCAI FLARE 2026 — joonasvirtanen · 2026-10-06
- Elad Gil: biology feels sleepy — last peak was CRISPR in 2012, AI should change that — dr_alphalyrae · 2026-10-06
- Prime Intellect presents verifier brittleness work at COLM, hiring across research — willcb · 2026-10-06