LlamaIndex benchmark scoring bug fixed, Datalab jumps from 65% to 93.6%

VikParuchuri found a flaw in LlamaIndex's benchmark scoring logic: the scorer penalizes extra fields in model output but fails to strip metadata/confidence fields before scoring, systematically underscoring API models—after the fix, Datalab's score jumped from 65% to 93.6%. He also noted that LlamaExtract Agentic Plus still tops the leaderboard after the fix, though its lead is questionable, and he publicly criticized how such basic errors get packaged into press releases and marketing materials, urging the industry to run its own evaluations.

Confirmed

Not Yet Confirmed

Why It Matters

2026-08-15 ~ 2026-08-15 · 6 related posts

Primary sources

1 near-duplicate retellings: VikParuchuri