JFPuget warns private benchmarks can leak into LLM training without no-retention terms
JFPuget · x · 2026-07-23
JFPuget argues that private benchmark data is not really private unless a strict no-retention contract is in place.
- The core concern is that prompts and benchmark inputs may be retained by the LLM provider after the first run.
- The post cites ARC-AGI-1 as a precedent, where OpenAI acknowledged using test data to train a preliminary version of o3.
- The implication is that benchmark leakage can happen even when the evaluation is meant to be private.
- The author’s bottom line: without hard contractual no-retention guarantees, private benchmarks can become training data.
More from Safety
- Quoted post says LLaMA 2 should be stopped to slow AI proliferation — 1a3orn · 2026-07-23
- AIRA proposal would regulate only frontier AI systems and require stronger security — 1a3orn · 2026-07-23
- U.S. lawmakers are preparing a bill that would let DHS throttle or shut down AI systems — The Verge AI · 2026-07-23
- HOL Guard adds local runtime protection for MCP servers and agent tool calls — kantorcodes1 · 2026-07-23
- Aidan Clark backs partnerships to safety-harden external models, not share the methods — _aidan_clark_ · 2026-07-23
- Researchers disclose one-click ChatGPT Workspace Agents hijack, fixed by OpenAI in 4 days — rez0__ · 2026-07-23