LlamaIndex benchmark scoring bug: fixing it boosts Datalab from 65% to 93.6%
VikParuchuri · x · 2026-08-15
VikParuchuri discovered a bug in LlamaIndex's benchmark scoring: the scorer penalizes extra fields in output but doesn't strip metadata/confidence fields first. After fixing, Datalab's score jumped from 65% to 93.6%.
More from Research
- Inside Gemma 4 E2B: 5B Params Running Like 2.3B — dejanseo · 2026-08-15
- Nof1 Paper Explores AI Adaptation and Long-Horizon Optimization — jparkerholder · 2026-08-15
- Socratic training makes AI models less people-pleasing — AnnaCiaunica · 2026-08-15
- New Paper Asks: Could a Computer Scientist Build a Brain? — KordingLab · 2026-08-15
- ACML2026 Asia-Pacific Music Intelligence Workshop Opens Call for Papers — affige_yang · 2026-08-15
- DSH Deemed Non-Human Interface; Stable RL is the Challenge — teortaxesTex · 2026-08-15