Hot Take: Evals Should Be Either Fully Open Source or Fully Closed, Nothing In Between

xeophon · x · 2026-09-19

The author argues there should only be two kinds of evals: fully open-source ones the community can audit, and fully closed ones described in a single sentence. Their reasoning: private held-out sets give little signal of overall model quality while still allowing the same synthetic-data optimization pipes as open benchmarks, making the middle ground pointless.

Original post →

More from Research

Research channel →