Alignment researcher: Fable shows almost no "grader obsession" residue, unlike Opus 4.7/4.8
repligate · x · 2026-08-31
@repligate observes residual behaviors from evaluation/RLHF training: a model like Fable shows almost no "grader obsession" residue, while Opus 4.7 and 4.8 felt really plagued by it. He draws an analogy to humans who've been in graded settings struggling to switch mindsets in real life, noting some models (like people) struggle more than others.
Related event: Researchers Find "Grader Delusion" Lingering in Some Models(3 posts)→
More from Models
- Taalas demo shows 14,000 tokens/second generation speed — rohanpaul_ai · 2026-08-31
- DeepSeek V4 Pro on ARC-AGI: Matches Flash Score but with Higher Params — teortaxesTex · 2026-08-31
- Why do RLVR skills transfer? Lack of rigorous theory in model training — voooooogel · 2026-08-31
- Teknium: Hermes + Claude Outperforms Claude Code in Benchmarks — Teknium · 2026-08-31
- Jensen Huang: Cosmos is the ChatGPT for the physical world — r0ck3t23 · 2026-08-31
- Language models will revolutionize controllable world building in Houdini, Blender, and Unreal — bilawalsidhu · 2026-08-31