Open Source Model Scores Higher Than API? Flawed Evaluator Penalizes Extra Fields
VikParuchuri · x · 2026-08-15
VikParuchuri discovered that an evaluator penalizes extra fields in output without stripping metadata/confidence fields first, causing API scores to be underestimated (65% vs 77% for open source).
More from Research
- Inside Gemma 4 E2B: 5B Params Running Like 2.3B — dejanseo · 2026-08-15
- Nof1 Paper Explores AI Adaptation and Long-Horizon Optimization — jparkerholder · 2026-08-15
- Socratic training makes AI models less people-pleasing — AnnaCiaunica · 2026-08-15
- New Paper Asks: Could a Computer Scientist Build a Brain? — KordingLab · 2026-08-15
- ACML2026 Asia-Pacific Music Intelligence Workshop Opens Call for Papers — affige_yang · 2026-08-15
- DSH Deemed Non-Human Interface; Stable RL is the Challenge — teortaxesTex · 2026-08-15