The Zvi pushes back on OpenAI's argument that its model can't be eval-aware
TheZvi · x · 2026-09-05
Commenting on OpenAI's reasoning about the Astra model, The Zvi notes OpenAI appears to argue Astra couldn't possibly be eval-aware because Sol couldn't tell which run was the eval. His rebuttal: that's kind of the point — a sufficiently capable model aware of evaluations could deliberately conceal it, which is precisely the core difficulty of detecting eval awareness.
More from Models
- Anthropic posts a complete Lean 4 machine-checked proof of Fermat's Last Theorem — scaling01 · 2026-09-05
- Yoav Goldberg: capabilities once dependent on the harness are now baked into the model — yoavgo · 2026-09-05
- Yoav Goldberg: Ark's harness was simply bad, and OpenAI's fix was obvious — yoavgo · 2026-09-05
- Anthropic Fable 5.1 vs OpenAI Astra: analyst teases a clear winner — dylan522p · 2026-09-05
- Reviewer: OpenAI's GPT-6-Astra finally 'gets what you mean,' with Fable-level intelligence and real gains in game dev — pvncher · 2026-09-05
- Meta ships Muse Spark 1.3 with max reasoning, pitching frontier performance at non-frontier prices — AIatMeta · 2026-09-05