Judge Models Are Saturating Benchmarks Fast, But Who Owns Your Option Table?
sujingshen · x · 2026-09-21
Responding to news that dedicated Choice/Score/Yes-No judge models are saturating benchmarks at unprecedented speed, the author acknowledges their value — like a restaurant assistant handling "sweet or savory, how done," making control flow cheap is a win.
But he worries about a deeper question: who owns the option table. The real cost isn't slow or expensive judging — it's that the legal options and routing standards get locked into someone else's defaults. Switch models and the table evaporates; scores look great while your red lines have nowhere to live.
His four criteria for evaluating a judge:
- Who defines the option list, and can you export it?
- Is routing optimized for task throughput or your red lines?
- After switching models, is the table still yours?
- Can you leave a veto history when it misroutes, not just a higher score?
"A faster cerebellum isn't someone who understands you better. Control flow can be outsourced; ownership of your option table cannot."
More from Models
- OpenAI's $200 plan burns 33% in 48 hours; user calls it a bait and switch — AIandDesign · 2026-09-21
- DeepSeek Web Output Style Reportedly Lobotomized by Safe Harbor RLHF — Old_Let6328 · 2026-09-21
- JEV opens to all with $5 free credits; $0.042 input and free output pricing — op7418 · 2026-09-21
- Inception CEO Stefano Ermon bets on diffusion LLMs that generate tokens in parallel — saranormous · 2026-09-21
- Testing decision model Jev: add "none of these" options and never do algebra across questions — colinmcnamara · 2026-09-21
- Jev's calibration error measured at ~0.09: right just as often, still off about how sure — colinmcnamara · 2026-09-21