Private evals under fire: hidden samples force people to guess how benchmarks work
xeophon · x · 2026-09-23
xeophon, quoting natolambert's "what is going on here" at a chart, argues the ugly graph says more about the eval than the models: with private evals you cannot inspect samples or runs, so people resort to weird theories about how it works — another jab at benchmark opacity.
More from Models
- Third-party test: Claude Opus 5.5 renders finer 3D scenes but costs 13x more than GPT-6 Sol — testingcatalog · 2026-09-23
- GPT-6 Sol priced at half of Opus 5.5 as Sol and Luna go 'dirt cheap' — ZeroStateReflex · 2026-09-23
- Tester claims Claude Opus 5.5 has the best visual design output of any model tested — burny_tech · 2026-09-23
- Meta's Alexandr Wang reveals muse has been in the works since at least Sept 2025 — adrianscottcom · 2026-09-23
- GPT-6 Sol Codex system prompt leaked: over 294,000 characters dumped on GitHub — gaganghotra_ · 2026-09-23
- Claude 5.5 (live) keeps generating user turns, reports user — BlackHC · 2026-09-23