LlamaExtract Agentic Plus still leads benchmark, but overfitting or scoring errors possible
VikParuchuri · x · 2026-08-15
VikParuchuri notes that LlamaExtract Agentic Plus still leads the benchmark, but possible reasons include its strength on this document mix, benchmark overfitting (vendors often lead their own benchmarks), or significant scoring errors. He previously mentioned that open-source model lift scored higher than API (77% vs 65%), suggesting scoring issues.
Related event: LlamaIndex benchmark scoring bug fixed, Datalab jumps from 65% to 93.6%(6 posts)→
More from Models
- Qwen 3.8 takes 10 mins for deep research vs Glimmer's 3, user says — Numerous_Mulberry514 · 2026-08-15
- Benchmark: Qwen3.8-27B medium effort matches 3.6 with 3x speed — peculiar-ragdoll · 2026-08-15
- Qwen3.8-27B slightly slower than predecessors on Apple Silicon — PerfectOlive1324 · 2026-08-15
- User suspects Qwen3.8-27B pruned general knowledge for coding skills — bonobomaster · 2026-08-15
- Qwen 3.8 27B Beats Claude Opus 4.6 in Three.js Coding Test for Free — testingcatalog · 2026-08-15
- Qwen3.8 27b dubbed 'new Microsoft Phi' in Reddit post — CavalryArcher · 2026-08-15