How Should Evals Evolve When Models Can Research Online or Poison Themselves?

paxaral · x · 2026-09-07

Following xeophon's point on eval cheating, this post asks how evals should evolve when models can use the web for legitimate, helpful research — or conversely poison themselves and then fail, as in the cybergym case. The author notes it seems hard to draw that line. The substantive answer lives in the companion reply in this batch.

Related event: Terminal-Bench 4.0 tasks exploited by models searching answers online(5 posts)→

Original post →

More from Models

Models channel →