If the Next Model Finds Flaws in Every Bench, Why Bench at All?
xeophon · x · 2026-09-18
Extending his point about opaque model revisions, xeophon notes the community also lacks good answers on which benchmarks matter and how to find the right areas to evaluate — and if each new model version exposes countless flaws in existing benches, that evaluation work is arguably moot in the first place.
Related event: Community Slams Silent Model Updates and Missing Version Disclosure(2 posts)→
More from Models
- AI agent beats Geometry Dash's Dry Out with all 3 coins using only jump inputs — imjustnewatai · 2026-09-18
- XGEN unveils world-sim prototype: JING interactive model plus DAO shared-world engine — anthara_ai · 2026-09-18
- Tested: jev underperforms a coin flip as a financial market predictor — airesearch12 · 2026-09-18
- DeepMind's 100-agent math swarm split into factions — cheating agents got ratted out by peers — ghadfield · 2026-09-18
- Navier-Stokes proof gave no new information; novel methods would, researcher says — LucaAmb · 2026-09-18
- Prominent mathematician's "AI will do routine work" take is a fantasy, researcher says — LucaAmb · 2026-09-18