JevBench details revealed: private question pools and sealed test sets to stop benchmark gaming
airesearch12 · x · 2026-09-30
The JevBench team responds to criticism by detailing its anti-gaming design:
- Mostly private pool: Most of the score comes from a private set of items no model builder ever sees; every release gets a fresh sealed set.
- Penalty for gaps: A model that scores clearly better on public items than sealed ones gets penalized for overfitting.
- Limited submissions: No unlimited submissions to the test set, by design.
The author also addresses trust in Homebrew/Linux dependencies: it's trust built on verification — open code, thousands of reviewers per change, SHA-256 checksums — not trust instead of verification. Both sides agree the goal is helping users pick the right model for their job.
More from Models
- OpenAI staffer: ChatGPT drives my computer better than I do — BorisMPower · 2026-09-30
- DevDay's best reveal is GPT-6.1 Sol, but dev says Astra would have been far more compelling — bindureddy · 2026-09-30
- OpenAI introduces Pro 500 plan: no 5-hour cap and 25x Plus usage — danshipper · 2026-09-30
- Early-access tester claims GPT-6.1 Sol matches GPT-6 Astra, faster and much cheaper — DeryaTR_ · 2026-09-30
- Reddit asks: is ChatGPT Dots just OpenClaw in a nicer suit? — Known-Beautiful-436 · 2026-09-30
- OpenAI offers zero data retention guarantee even at inference time for enterprises — BorisMPower · 2026-09-30