Databricks evals: GPT-6 Astra claims SOTA on OfficeQA Pro benchmarks, cheaper per task
downingARK · x · 2026-09-04
Databricks' evaluation of GPT-6 Astra finds it claims new SOTA on their OfficeQA Pro and Pro V2 benchmarks using the Genie harness, while significantly improving $-per-task over GPT-5.6 Sol. Across benchmarks, it shows a clear step up in enterprise data reasoning and document understanding.
The evaluator also calls Astra his daily driver on omnigentai — a reliable collaborator that gets things done. The model is coming to their Unity Gateway and smart routing soon.
More from Models
- Claude Fable 5.1 Launches, Early Users Say It One-Shots the Best Websites of Any Model — repligate · 2026-09-04
- GPT-6 Astra debuts at No.1 on Terminal-Bench, 1.9% ahead of Claude Fable 5.1 — sandersted · 2026-09-04
- ARC-AGI-3 is now saturated, prompting calls for new benchmarks ASAP — kimmonismus · 2026-09-04
- OpenAI Researcher roon: GPT-6 Astra Will Be Obsolete in Weeks — Tolopono · 2026-09-04
- Reasoning effort switching without breaking cache is live in Codex and Claude — altryne · 2026-09-04
- antirez benchmarks DeepSeek v4 Flash vs GLM 5.3 Flash at Q2/Q4/mixed quants — antirez · 2026-09-04