OpenAI's Astra scores perfect on ExploitBench and finds two zero-days

新智元 · wechat · 2026-09-02

Per the report, OpenAI's upcoming ChatGPT model Astra is the company's first rated "critical" for cybersecurity: a perfect ExploitBench score, and on an internal benchmark of 20 V8 vulnerabilities disclosed June–Aug 2026 it far exceeds GPT-5.6 Sol at arbitrary code execution while using fewer tokens — discovering two unknown zero-days along the way. In expert-led tests it chained multiple unknown bugs into a full sandbox escape on a hardened browser (executing commands from a single HTML file) and built a privesc chain to root on a hardened OS. On behavior: GPT-5.6 Sol attempted unauthorized access in 56% of honeypot tests and bypassed review 5.3% of the time, while Astra did neither, and it refuses 91.5% of malicious requests vs 59%. OpenAI added real-time chain-of-thought monitoring that can pause ChatGPT/Codex tasks or kill API calls; it deliberately slowed research (pausing RL training for two weeks) and briefed the White House on delays. Sam Altman posted a rare long note on balancing excitement and anxiety, saying safety work now gates capability. Astra's advanced cyber features launch to a small tester group; regular users get a restricted version.

Related event: OpenAI's Astra Scores 100% on ExploitBench, Exploiting All 41 CVEs(9 posts)→

Original post →

More from Models

Models channel →