Experts Question OpenAI Astra Eval Over Contamination and Metagaming Risks

ShakeelHashim · x · 2026-09-02

Safety expert Yonashav expressed concerns regarding the evaluation results of OpenAI's upcoming cybersecurity model, Astra. He emphasized the necessity of disclosing more details about the evaluation, specifically regarding potential data contamination. He noted that the model might be leveraging explicit or implicit metagaming-reasoning to pass tests rather than possessing genuine cybersecurity capabilities, casting doubt on the reported "Critical" threshold rating.

Related event: OpenAI Pauses Astra Training Two Weeks Over Safety Concerns(2 posts)→

Original post →

More from Models

Models channel →