A regex script scores 95% on User Sim Index, researcher warns eval is broken
ericzelikman · x · 2026-10-06
Researcher Eric Zelikman shows the popular "User Sim Index" for evaluating user models is gameable: with attached code, a trivial rule-based "user model" scores 97% comms, 97% info, 92% clarify, and 95% error reaction — beating any released model on those dimensions.
He clarifies that regex/statistical signals can have value, but such results shouldn't be presented as headline findings. A cautionary note on the reliability of behavioral eval benchmarks.
Related event: Researcher Shows User Sim Index Benchmark Easily Gamed to 95%(3 posts)→
More from Models
- AI eval researcher: even humans can't detect subtle AI patterns like distributional biases — alexisjross · 2026-10-06
- At $20/month, OpenAI gives you everything, Google quietly tiers models by surface — bytebot · 2026-10-06
- ChatGPT puts real cartoonists' signatures on fake New Yorker cartoons — luisdans · 2026-10-06
- Dev bets Thinking Machines' "fledge alpha" will ship as "fledgling" — willcb · 2026-10-06
- Insider speculation: GPT-6.5 expected within 4 weeks, above Fable 5.5 level — VraserX · 2026-10-06
- Gemini users hit image generation limits as Google tightens daily quotas — BrattyMiku · 2026-10-06