Eval anti-cheat idea: serve models a stale HF cache from before grader fixes
willcb · x · 2026-08-28
In a short exchange on eval contamination, willcb suggests giving models a stale HuggingFace cache from before the eval author inspected the data and patched grader bugs; the reply argues it'd be better to just turn off web access instead of praying. A quick debate on preventing benchmark leakage and cheating.
More from coding & agent
- Developer Publishes Handbook on RAG and Context Engineering Based on Real Papers — techNmak · 2026-08-28
- Claude Autonomously Ships to SaaS: Ring-Based Permission System — mhmazur · 2026-08-28
- Grok Bot Guides: Practical Playbooks for Multi-Agent Collaboration — XFreeze · 2026-08-28
- Complete AWS Agent Deployment Guide: From Terraform to CI/CD — blaizedsouza · 2026-08-28
- Claude + Thrixel Skills build alchemist Water Sort game in ~1 hour — RanaHanocka · 2026-08-28
- OtelJazz: Sonifying multi-agent telemetry to make drift and failures audible — mob1ius · 2026-08-28