User Sim Index is broken: trivial bot scores 95% across behavioral dims
ericzelikman · x · 2026-10-06
ericzelikman argues the community should stop evaluating user models on the "User Sim Index": he shares code for a simple "user model" scoring 95% overall on 4 behavioral dims (97% comms, 97% info, 92% clarify, 95% error reaction), claiming it beats any released model — evidence the benchmark is gameable and not measuring real user-simulation quality.
Related event: Researcher Shows User Sim Index Benchmark Easily Gamed to 95%(3 posts)→
More from Research
- Is the brain a computer? Physical reservoir computing pokes holes in Turing-simulation argument — cephaloform · 2026-10-06
- Why User-Model Evals Are Hard: Stanford Researchers Bet on a Multi-User Turing Test — alexisjross · 2026-10-06
- Embedding Every Font with Neural Networks Yields a Flower-Shaped Map of Google Fonts — Chroma-Crash · 2026-10-06
- Crawler Zoo Launches a Free Arena for Testing Local-Model Agents — Time_Instruction_955 · 2026-10-06
- Trained agentic context management: 8K-context small model matches GPT-5.4 at 1M on OOLONG — xennygrimmato_ · 2026-10-06
- Used OpenAI Dots as a Free Agent Swarm to Break a 47-Year-Old Math Record — jaxchang · 2026-10-06