JevBench Creator: Picking Models by 'Vibes and Trust' Is a Bad Joke
airesearch12 · x · 2026-09-29
JevBench creator @airesearch12 pushes back on @CompleteSkeptic's claim that public benchmarks miss the point and are trivially gameable, arguing that picking a System One model by "vibes and trust" is a bad joke or a way of coping with the lack of a moat in model development:
- Benchmarks aren't perfect, but they remain the best objective, scientific proxy for model quality.
- JevBench is deliberately built so scores can't be "benchmaxxed": metrics and test-set composition change every release, so model builders can't train against them.
- His benchmark site also offers what he calls the world's first benchmaxxing score.
At heart it's the debate between public evaluation and subjective feel.
More from AGI Musings
- David Patterson: humanoid robots will tunnel cities' streets underground into parks — davidpattersonx · 2026-09-29
- Dan Jeffries: RSI Will Hit a Hidden S-Curve, Leaving AI Useful but Not Magical — Dan_Jeffries1 · 2026-09-29
- Are smarter people less likely to cheat? An AI researcher's observation and hope for AI — gandamu_ml · 2026-09-29
- David Patterson: AGI will replace knowledge and physical work at the same time — davidpattersonx · 2026-09-29
- Blogger Proposes RLEI, a Replacement for RLVR Based on Epistemic Incompleteness — ryunuck · 2026-09-29
- As Models Surpass Human IQ, Alignment May Mean Answers Humans Understand — djcows · 2026-09-29