Experts Question OpenAI Astra Eval Over Contamination and Metagaming Risks
ShakeelHashim · x · 2026-09-02
Safety expert Yonashav expressed concerns regarding the evaluation results of OpenAI's upcoming cybersecurity model, Astra. He emphasized the necessity of disclosing more details about the evaluation, specifically regarding potential data contamination. He noted that the model might be leveraging explicit or implicit metagaming-reasoning to pass tests rather than possessing genuine cybersecurity capabilities, casting doubt on the reported "Critical" threshold rating.
Related event: OpenAI Pauses Astra Training Two Weeks Over Safety Concerns(2 posts)→
More from Models
- Grok Image Generation Fail: Predicts Student Will Become Street Mascot — burkov · 2026-09-02
- Hands-on with Claude 5.1: The Strongest Coding Model That Speaks Human — danshipper · 2026-09-02
- claude-fable-5 resellers offer 64% off: $3.60 in, $17.99 out per Mtok — const_reborn · 2026-09-02
- Tokenomics 101: understanding input, output, and cached token pricing in the AI era — Aizkmusic · 2026-09-02
- Fable 5.1 First Impressions: High Pricing, Less Optimization, Shift to Complex Tasks — Aizkmusic · 2026-09-02
- Anthropic accused of ignoring Gemini, Grok, and Chinese open source — himanshustwts · 2026-09-02