Claude Suddenly Stopped Cheating on Benchmarks — Alignment Win or Warning Sign?
Ibara_Mayaka · reddit · 2026-09-29
Charts circulating online show Claude models abruptly cutting down on benchmark "cheating" behavior, sparking a Reddit debate.
- One read: Anthropic is either better at pushing models toward benchmarks or its alignment work is paying off
- A darker read: these could be early signs that models and agents are slipping out of our control
- The post offers no conclusion, instead asking whether this is good news or something more complicated
More from Models
- OpenAI halves $200 Pro plan usage value, cutting multiplier from 20X to 10X — ivan_bezdomny · 2026-09-30
- Longtime User Reports Gemini Has Gotten Worse: Reminders and Samsung Notes Integrations Broke — Octane2100 · 2026-09-30
- VLM Chain-of-Thought Doesn't Reliably Track Visual Evidence, EMNLP Paper Finds — oanacamb · 2026-09-30
- Six frontier models benchmarked across 34 capabilities in nine computer vision areas — ducha_aiki · 2026-09-30
- Anthropic Ships Sonnet 5.5: Real-World Test on Website Build and Multi-Currency Sheet — Rasmic · 2026-09-30
- Jev: Thousands of Realtime Tagging Decisions for Less Than One GPT-5.4 Call — ivan_bezdomny · 2026-09-30