OpenAI’s Codex + GPT-5.6 Sol hits 99% recall in Project APE verification tests
soumitrashukla9 · x · 2026-07-22
The first key result from the Project APE work is a major jump in automated verification performance.
- Just six months ago, the best verifier was below 80% and missed 1 in 5 real errors.
- In July 2026, the team says it reached 99% detection recall for the first time.
- The model cited for that result was Codex + GPT-5.6 Sol from OpenAI.
- Days later, Kimi K3 reportedly reached 98%.
More from Models
- Security Differences Between Closed and Open Source Models: Insights from OpenAI's Escape Incident — robleclerc · 2026-07-23
- DeepSeek V4 and Kimi K3 Announced as Imminent Amidst AI Acceleration — emmanuelvivier · 2026-07-23
- Google Reportedly Starts Gemini 4 Pre-training in Most Ambitious Run Yet — emmanuelvivier · 2026-07-23
- Google Launches 3 New Gemini Models: 3.6 Flash Cuts Costs and Output Tokens — emmanuelvivier · 2026-07-23
- Meme compares Gemini 4’s progress to GPT-5.4 mini’s lead — cgarciae88 · 2026-07-23
- Google Gemini reportedly reaches 950M monthly users and 22B API tokens a minute — zephyr_z9 · 2026-07-23