OpenJev Benchmark Reworked: Accuracy Weighted Higher, Probability Calibration Added
airesearch12 · x · 2026-09-19
The openjev benchmark author revised his earlier equal-weight scheme, tilting scoring toward accuracy since a dirt-cheap but wildly inaccurate model is trivial to build. He also plans to add a new axis: accuracy on probabilities, i.e. calibration. The design is still a work in progress.
Related event: Nearly 20 openjev models emerge as enthusiast launches leaderboard(4 posts)→
More from Research
- Sebastian Raschka Releases Inference Scaling Tutorial: Self-Consistency Boosts Accuracy Over 2x — rasbt · 2026-09-19
- AI interpretability community still fighting 'SAE is dead' claims, one airport at a time — EigenGender · 2026-09-19
- Author's submission ID 62263 suggests ICLR 2027 total near 63K papers — trawasthi_ai · 2026-09-19
- ICLR 2027 hits 62K+ submissions, peer review faces its biggest stress test yet — sivareddyg · 2026-09-19
- Sebastian Raschka's inference scaling video: building self-consistency that lifts accuracy >2x — rasbt · 2026-09-19
- WeightWatcher experiments suggest Muon beats AdamW by reducing memorization — rasbt · 2026-09-19