Databricks evals: GPT-6 Astra claims SOTA on OfficeQA Pro benchmarks, cheaper per task

downingARK · x · 2026-09-04

Databricks' evaluation of GPT-6 Astra finds it claims new SOTA on their OfficeQA Pro and Pro V2 benchmarks using the Genie harness, while significantly improving $-per-task over GPT-5.6 Sol. Across benchmarks, it shows a clear step up in enterprise data reasoning and document understanding.

The evaluator also calls Astra his daily driver on omnigentai — a reliable collaborator that gets things done. The model is coming to their Unity Gateway and smart routing soon.

Original post →

More from Models

Models channel →