giffmana skeptical: found training env already contaminated, eval protections unlikely to hold

giffmana · x · 2026-09-11

In a thread about benchmark contamination, MaxKannen argues the discovered environment may only be used for training, not eval, and notes Anthropic claims to invest in making eval environments unidentifiable. giffmana responds that even if this one random environment was training-only, its very discovery shows how easy contamination is — "we will never really know" — and he isn't holding his breath for the protections to work.

Related event: Debate Over Eval Environment Contamination as Anthropic Defenses Questioned(3 posts)→

Original post →

More from Models

Models channel →