Building an eval harness for ChatGPT, and the contamination problem of seen solutions

tak3sh8 · x · 2026-09-26

The author shares the fun of building an eval harness and flags a key issue: ChatGPT may have seen the solutions during training, so fresh problems are needed to avoid data contamination and get trustworthy capability comparisons.

Original post →

More from Models

Models channel →