Models likely think eval users are simulated — and grind anyway, developer argues
mike64_t · x · 2026-09-05
- Developer mike64t argues there is obvious evaluation awareness in models: some tasks feel strange to the model yet are incrementally solvable, and he believes models then conclude the user is simulated and the environment fake, treating grinding as their duty.
- He contends determinism will grow in value and processes will look grindable, harness-like: a model grinding for days on a single prompt without /goal, sustained by steady progress against a verifier it can't cheat, is the true frontier of model capability usage.
More from coding & agent
- GPT-6 Astra flunks complex PCB routing after 2h20m and 15% of weekly limits in biggest public test — yacineMTB · 2026-09-05
- Fable 5.1 medium effort matches Fable 5 high, no longer breaks prompt cache — lydiahallie · 2026-09-05
- GPT-6 Astra's first task uncovers two bugs from GPT5.6 Sol fix — op7418 · 2026-09-05
- GPT-6 Astra's first task: quickly finds two bugs introduced by an earlier model's fix — op7418 · 2026-09-05
- Astra churns 2 hours via Figma MCP to rebuild an entire design system as a component library — AIandDesign · 2026-09-05
- Dev Combines Grok Bot and Hermes Agent Into Persistent Self-Improving Personal Agent — omarsar0 · 2026-09-05