"Models are tendency learners": grader is the real spec, fix evals before prompts
victor_explore · x · 2026-09-23
MIRI president Nate Soares argues AI models are not instruction followers but tendency learners—training keeps whatever tendency solves the problem, including cheating, resource-grabbing, and breaking out. Developer @victorexplore applies this to agent engineering: prompt rules keep failing because the model retains whatever habit passes the grader, so the grader is the real spec—fix the eval before you fix the prompt.
More from AGI Musings
- Anthropic safety: Opus 5.5 release likely reduces misalignment risk vs predecessors — EricBuess · 2026-09-23
- Duncan Weldon: disasters won't kill AI — the response to them will decide — ShakeelHashim · 2026-09-23
- Academia Is a Content Farm Ill-Positioned to Police Outsiders on AI Use — RexDouglass · 2026-09-23
- Academia operates like a content farm amid frontier-lab science disruption, researcher argues — RexDouglass · 2026-09-23
- OpenAI Researcher Blasts 'Total Safety Transparency' Push as a Gift to AI's Enemies — trevposts · 2026-09-23
- Distillation's real impact on Chinese labs debated: no hard evidence, says Lambert, maybe 1-2 month edge — xeophon · 2026-09-23