Perplexity's WANDR Eval: GPT-6 Astra Tops at 0.682, $11.98 Per Task
cameronstow · x · 2026-09-04
Perplexity published its evaluation of GPT-6 Astra on its in-house benchmark WANDR: the model scored 0.682 — the highest of any model they tested — at $11.98 per task.
Comparisons: GPT-6-Astra scored 13.5% higher than Fable 5.1 at 6.1% lower cost, and 27.0% higher than Opus 5 at just 3.3% higher cost. A systematic third-party capability-vs-price evaluation useful for model selection.
More from Models
- OpenAI rolls out misalignment monitoring for Astra, admits it may miss harmful behavior — ShakeelHashim · 2026-09-04
- Databricks evals: GPT-6 Astra claims SOTA on OfficeQA Pro benchmarks, cheaper per task — downingARK · 2026-09-04
- Mathematician tests GPT-6 Astra: live Lean proof verification while writing arguments — teortaxesTex · 2026-09-04
- Tavus Launches Sparrow-2, Claiming #1 in End-of-Turn Detection and Interruption Handling — ycombinator · 2026-09-04
- ARC-AGI-3: 10x reasoning tokens cuts total cost from $48k to $26k vs medium — i_dg23 · 2026-09-04
- Researchers flag data contamination concerns in benchmark behind Astra's time-horizon score — dfrsrchtwts · 2026-09-04