Hot Take: Evals Should Be Either Fully Open Source or Fully Closed, Nothing In Between
xeophon · x · 2026-09-19
The author argues there should only be two kinds of evals: fully open-source ones the community can audit, and fully closed ones described in a single sentence. Their reasoning: private held-out sets give little signal of overall model quality while still allowing the same synthetic-data optimization pipes as open benchmarks, making the middle ground pointless.
More from Research
- JevBench v1 puts nine typed-decision models head-to-head across 242 decisions — airesearch12 · 2026-09-19
- FlashREINFORCE debuts: critic-free, single-rollout async RL for agentic LLMs — CatAstro_Piyush · 2026-09-19
- AI companies are conquering math — and exposing a discipline built on competition, not understanding — danbri · 2026-09-19
- CMU Autonomous Science Lab to Feature at Enamine Drug Discovery Conference — olexandr · 2026-09-19
- AI & Science feature in Scientific American gets a positive write-up — JMateosGarcia · 2026-09-19
- Terence Tao: If Math Is More Than Proof, We Must Celebrate the Rest of It — num42 · 2026-09-19