GPT-6 Astra hits 97.2% on short-doc extraction SOTA but only 31.7% on long docs
llama_index · x · 2026-09-06
LlamaIndex (Jerry Liu) benchmarked GPT-6 Astra on hard document parsing and extraction:
- 97.2% one-shot extraction on short documents, a new SOTA on their ExtractBench; 90.6% on medium documents
- Long documents remain hard: only 31.7% one-shot
- Pricey: 11c per page, 10x their cost-effective extraction solution
- Native OCR is not much better than peers like Fable 5.1 or gpt-5.6-sol — strong on tables, weak on charts, layout, and semantic formatting
Bottom line: best frontier model for short/medium document extraction, but cost and long-doc performance are real caveats.
More from Models
- Astra Benchmark Performance Still Poor, Says Early Tester — gleech · 2026-09-06
- LLMs Play Chess Poorly, So This One Built Its Own Chess Engine to Fight Back — MikePFrank · 2026-09-06
- Astra Computer Access Fails Drag-and-Drop Web Page Build Test — BLUECOW009 · 2026-09-06
- Astra disappoints on harness-building tasks while Fable 5.1 excels, dev reports — HarveenChadha · 2026-09-06
- GPT-6 Astra runs DOOM at 20+ FPS on its own CPU, up from GPT-5.6 Sol's 1 FPS — Angaisb_ · 2026-09-06
- GPT-6 Astra designs a working jet plant and ships a live 3D simulation autonomously — deanwball · 2026-09-06