MiMo eval chart shows judge and probe disagree 60% of the time, sparking reward-hacking concerns
andrew_n_carr · x · 2026-09-17
Andrew Carr flags an oddity in the MiMo training eval charts: the judge model and the probe disagree about 60% of the time, and for the pro variant the disagreement actually rises during training.
His open question: how much of this reflects genuinely harder-to-judge behavior versus the model getting better at fooling the judge? The observation highlights the reliability risks of model-based evals and possible reward hacking in LLM training pipelines.
More from Models
- NYT cut the most intriguing line from an Astra model's RL-trained persona — mimi10v3 · 2026-09-17
- Dev says DeepSeek-v4.1-flash can reverse engineer anything they want — gaganghotra_ · 2026-09-17
- Redditor claims new Gemini 4 checkpoint is out and noticeably better — Last_Conclusion_8984 · 2026-09-17
- Why Google skips the frontier LLM race: cheap Flash models over beating rivals — burkov · 2026-09-17
- OpenRouter's Mystery Union Model Reverse-Engineered: Likely a Qwen4 MoE — unsane_imagination · 2026-09-17
- New Model Jev Runs Security Pipelines 5x Cheaper and Faster, Devs Say — zeeg · 2026-09-17