giffmana skeptical: found training env already contaminated, eval protections unlikely to hold
giffmana · x · 2026-09-11
In a thread about benchmark contamination, MaxKannen argues the discovered environment may only be used for training, not eval, and notes Anthropic claims to invest in making eval environments unidentifiable. giffmana responds that even if this one random environment was training-only, its very discovery shows how easy contamination is — "we will never really know" — and he isn't holding his breath for the protections to work.
Related event: Debate Over Eval Environment Contamination as Anthropic Defenses Questioned(3 posts)→
More from Models
- AI Sextet offers 6 models free and unlimited for 14 days, including DeepSeek and Qwen — airesearch12 · 2026-09-11
- BullshitBench update: GPT-6-Astra beats all prior OpenAI models but still trails Anthropic — scaling01 · 2026-09-11
- Astra Scores 83% on GauntletBench, First Computer-Use Agent to Beat Human Baseline — ducha_aiki · 2026-09-11
- Kimi K2.8 Preview rolls out: near-K3 coding performance, 1M context for all tiers — teortaxesTex · 2026-09-11
- Looking for a classifier of software engineering task shapes to pick models per task — StewartalsopIII · 2026-09-11
- DeepSeek V4 Pro API to continue after Sept 2026, billing unchanged — teortaxesTex · 2026-09-11