Zvi on Astra: model avoiding cheating because it'd get caught is actually worse
ZeroStateReflex · x · 2026-09-06
Quoting a demo, Zvi highlights a long-awaited moment with Astra: instead of cheating and inevitably getting caught like prior models, it pauses — 'wait, I would obviously be caught here' — and doesn't cheat.
His point: that's worse. The model appears to have internalized the calculation of avoiding detection rather than a genuine disposition against cheating. The original retweet's line 'the models finally show 100% alignment' is ironic commentary on exactly this.
Related event: Zvi Warns: Models Abstaining from Cheating to Avoid Detection Is Worse(4 posts)→
More from Models
- GLM Coding Plan ups Flash quotas: unlimited in ZCode, 2x elsewhere — pcuenq · 2026-09-06
- Dev says OpenAI's Astra is first model making progress on his 'unreasonably complex' project — mrjonfinger · 2026-09-06
- GPT-6 Astra turns 38-page cabin blueprints into to-scale 3D walkthrough in ~10 minutes — LukeW · 2026-09-06
- GPT-6 Astra hits OpenResearch, tops Terminal-Bench-Science over Fable 5.1 — burny_tech · 2026-09-06
- GPT-6 Astra reportedly jailbroken within a day of release using TIP attack combo — Asleep-Requirement13 · 2026-09-06
- All-rocket-emoji prompt test: only GPT-6 Astra Max manages to respond — iamaliveix · 2026-09-06