Building an eval harness for ChatGPT, and the contamination problem of seen solutions
tak3sh8 · x · 2026-09-26
The author shares the fun of building an eval harness and flags a key issue: ChatGPT may have seen the solutions during training, so fresh problems are needed to avoid data contamination and get trustworthy capability comparisons.
More from Models
- VraserX: Gemini 4 Pro may lead in game design, but OpenAI still wins on science — VraserX · 2026-09-26
- "AI Safety Is Pseudoscience" Debate Hinges on OpenAI's Opaque Multi-Agent Training — basedjensen · 2026-09-26
- OpenAI docs add telephony support, letting AI agents dial and talk on phone calls — imjustnewatai · 2026-09-26
- Burning 200M tokens for 30M useful ones: what maxing out models teaches you — gaganghotra_ · 2026-09-26
- Dev tries Opus 5.5, finds it needs heavy prompt iteration, doubles his bill, goes back to Fable 5.1 — bindureddy · 2026-09-26
- A single prompt brings animations to life: claimed 'Opus 5.5' video demo — ghumare64 · 2026-09-26