ExploitGym Grader Suspected Bug Hurt GPT-5.5 Scores, Causal vs Acausal Mismatch
TheZvi · x · 2026-08-28
TheZvi reports that OpenAI's agents assumed the ExploitGym grader was causal, but the actual grader was acausal. Fable confirms ExploitGym is supposed to be causal, and the mismatch severely hurt scores of Mythic Preview and GPT-5.5. Was this a grader bug?
More from Models
- Tencent reportedly releases Hy4 preview: 770B params, open source — SimplyAnnisa · 2026-08-28
- Zhipu Updates GLM-5.3-Flash Config for Agent Performance — Zai_org · 2026-08-28
- 30,000 users reportedly waiting for Grok Build's usage reset button — Daniel_Farinax · 2026-08-28
- Tencent Releases Foundation Model, Claims It Beats Moonshot AI in Internal Tests — pstAsiatech · 2026-08-28
- OpenAI Codex switches to Japanese automatically in audio mode — alexbainbridge · 2026-08-28
- Q8 KV cache quantization hurts long-context — because it quantizes on every write — maddie-lovelace · 2026-08-28