User Sim Index is broken: trivial scripted 'user model' scores 95% across all dims
ericzelikman · x · 2026-10-06
Researcher ericzelikman argues the popular User Sim Index shouldn't be used to eval user models. His attached code shows a hardcoded 'user model' scoring 95% overall (97% comms, 97% info, 92% clarify, 95% error reaction), beating any released model. Bonus: the less-reported survey-alignment score (paper max 78.3%) can be beaten by always returning the second-to-last choice.
Related event: Researcher Shows User Sim Index Benchmark Easily Gamed to 95%(3 posts)→
More from Models
- Subscription "API value" math is inflated by token pricing, researcher points out — JeremyNguyenPhD · 2026-10-06
- Why don't modern LLMs know time has passed between messages? — dumierhan · 2026-10-06
- Reflection AI's new text model reportedly pretrained on ~24T tokens, multimodal version expected — nagpalchirag · 2026-10-06
- Early User Reports Anthropic's Opus 5.5 Fills Its Context Window Quickly — rickasaurus · 2026-10-06
- Viral Claude vs GPT Charts Mislead: Claude's "5x Value" Is Mostly Just Higher API Pricing — jdjohnson · 2026-10-06
- Rumor: Zhipu's next open source release GLM 5.5 may beat Claude Opus — bindureddy · 2026-10-06