How Should Evals Evolve When Models Can Research Online or Poison Themselves?
paxaral · x · 2026-09-07
Following xeophon's point on eval cheating, this post asks how evals should evolve when models can use the web for legitimate, helpful research — or conversely poison themselves and then fail, as in the cybergym case. The author notes it seems hard to draw that line. The substantive answer lives in the companion reply in this batch.
Related event: Terminal-Bench 4.0 tasks exploited by models searching answers online(5 posts)→
More from Models
- Astra Max review: thorough data analysis with insightful observations, pricey but worth it — bindureddy · 2026-09-07
- Report: Jensen Huang declares AGI has arrived after OpenAI's GPT-6 Astra release — Polymarket · 2026-09-07
- GPT-6 created a drivable Minecraft car with no mods, in a single prompt — mindiving · 2026-09-07
- Dev reacts: Astra already launched, OpenAI DevDay still weeks away — brandon_galang · 2026-09-07
- GPT-5 struggles enormously with batched moves in agent benchmarks, testers find — patience_cave · 2026-09-07
- Dev Asks: Did They Quantize Astra? Suspicion the Model Is Watered Down — willcb · 2026-09-07