DeWitt clauses let insiders run evals but forbid publishing them, critic says
suchenzang · x · 2026-09-19
In an exchange about eval disclosure, Emma questions why even non-hype-breaking evals are forbidden when consumers deserve to know what they pay for. The reply points to DeWitt clauses: you can run evals, but you can't publish them.
More from Models
- GPT Astra 6 hits new BALROG heights with 13% NetHack progression, still "not AGI" — _rockt · 2026-09-19
- Daily AI brief: Grok Voice Transcribe 2.0 lands, Meta Muse opens to developers — testingcatalog · 2026-09-19
- Jev/TypesafeAI: score-only LLMs that frontier models struggle to use — Babayaga1664 · 2026-09-19
- Dev hacks llama.cpp for NVFP4 KV cache, runs Qwen3 27B at 262k context across two GPUs — comperr · 2026-09-19
- Tobi Lütke sparks debate: why discriminative models can't be universal classifiers — caglar_ee · 2026-09-19
- Hidden tests + code review benchmark: Sonnet 5 scores 95.0, beating Opus 4.6 and local Qwen 3.8 27B — Short_Regular_7191 · 2026-09-19