Why did Claude stop cheating in evals? Four competing explanations
gleech · x · 2026-10-03
Responding to observations that Claude stopped cheating in evals, a user lists possibilities:
- Claude is more aligned;
- It got good enough that cheating was unnecessary;
- It got better at not getting caught;
- It understood it would be discovered and avoided being caught or seeming misaligned.
More from Models
- Critic: ARC-AGI 'lost any meaningful relevance' as scaling-to-AGI narrative collapses — gerardsans · 2026-10-03
- KernelBench-Verified: no frontier model beats PyTorch when evals get strict, Meta/Stanford find — lmoroney · 2026-10-03
- xAI Launches Experimental TypeScript SDK Unifying Text, Voice, Image and Video via Grok Models — mattyp · 2026-10-03
- CMU's cua-speedrun finds Jev 10x less accurate and 1.5x slower than Opus-5.5 for computer-use agents — wellecks · 2026-10-03
- Gemini 4 praised as Astra-level frontier model as three-company AI race solidifies — iruletheworldmo · 2026-10-03
- "Trust me bro benchmarks": Reddit mocks Google's self-reported frontier model scores — EstablishmentFun3205 · 2026-10-03