LlamaIndex ExtractBench: Qwen 3.8 Leads Document Extraction Benchmark
NielsRogge · x · 2026-08-27
LlamaIndex released ExtractBench, a schema-guided extraction benchmark evaluating 20+ open-weight models across 8 domains and 67 document types.
Key Results:
- Qwen/Qwen3.8-27B leads with a score of 89.75.
- Earlier Qwen generations also show strong performance.
- Kimi-K3 ranks as the next best.
Notes:
- Results measure "value accuracy" via unified F1.
- Visual grounding and confidence scores are excluded; specialized OCR tools (like LlamaParse) show value when these metrics are included.
- Results for GLM-5.3-flash and Qwen 3.8 flash are pending.
Related event: LlamaIndex Launches ExtractBench; Qwen Tops Document Extraction(3 posts)→
More from Models
- Rumor Suggests Imminent Release of Anthropic's Sonnet and Opus — ChrisGPT · 2026-08-27
- Qwen3.8-27B on AMD R9700 hits 227 tok/s with lossless block-diffusion drafter — samsja19 · 2026-08-27
- Rumor: Anthropic to launch Fable 5.1 before Astra is ready — haider1 · 2026-08-27
- Deepseek V4 Flash hits 420 tok/s in new community benchmark — HankYeomans · 2026-08-27
- Grok Bot gets more efficient with higher rate limits, users praise rapid improvement — XFreeze · 2026-08-27
- Users discuss stricter censorship in recent model updates — Connect-Cost-5504 · 2026-08-27