Third-party AI evals hit a sharp ceiling if training info stays commercially secret
1a3orn · x · 2026-09-13
In a debate with @ThomWolf, AI researcher 1a3orn argues that third-party evals solve the principal-agent problem of grading your own homework, but that's not the same as advancing knowledge. If evaluators can't release training info because it's commercially sensitive, the benefit hits a sharp ceiling — public info is necessary for public science, and what's really needed is scrutiny from many people.
More from AGI Musings
- Garry Tan: become a domain-specific harness or die a system of record — AccBalanced · 2026-09-13
- 'Pacing the frontier' isn't stopping AI — the speed of progress itself is becoming the risk — flavioAd · 2026-09-13
- Dario's 'We Must Pace the Frontier' essay: Anthropic opens employee-level access to third-party evaluators — IgorCarron · 2026-09-13
- Instinct's agent UX is prized at $20-50B, but its relay-farm privacy design is a structural flaw — AccBalanced · 2026-09-13
- Counterpoint to AI pacing: whoever doesn't slow down becomes the new frontier — McDonaghMatthew · 2026-09-13
- Debate: 95% of tasks are doable with low-end or open-source models, so frontier hype is over — Sellao93 · 2026-09-13