Commenter warns a model's 3-hour benchmark data may be heavily cherry-picked
RexDouglass · x · 2026-10-09
- A commenter argues that without full metric disclosure, a model's widely shared "3-hour" evaluation result is likely cherry-picked.
- Possible tactics: running each problem 100 times and reporting only successes, or retrying across different model versions and publishing the best runs.
- The takeaway: treat evaluation claims lacking methodology transparency with skepticism.
More from Models
- Latest GPT is the reverse of "biased for action", users complain — altryne · 2026-10-09
- Codex throttled to 5 tok/s as dev argues local model deployment is the only fix — lxfater · 2026-10-09
- OpenAI's math claims criticized for skipping peer review, inverting the proper process — JFPuget · 2026-10-09
- Users poke fun as Anthropic tightens guardrails on swearing at its AI — CtrlAltDwayne · 2026-10-09
- Per-token cost of Claude Pro's Opus subscription works out the same as DeepSeek Flash — teortaxesTex · 2026-10-09
- DeepSeek's token usage on OpenRouter topped OpenAI, Google, Anthropic and xAI combined — FinanceYF5 · 2026-10-09