GPT-6 Astra tops Perplexity's WANDR benchmark at 0.682, 13.5% above Fable 5.1 at 6.1% lower cost
rohanpaul_ai · x · 2026-09-05
Perplexity published its evaluation of GPT-6 Astra on the WANDR benchmark, scoring 0.682 at $11.98 per task — the highest of any model tested.
Key numbers:
- 13.5% above Fable 5.1 while costing 6.1% less
- 27.0% above Opus 5 at just 3.3% higher cost
WANDR is unusual in testing wide-and-deep research: finding large sets of qualifying entities and backing every fact with checkable evidence. It contains 500 public tasks requiring 170,495 source-backed records, penalizing incomplete research even when found facts are correct, with strict hard scores requiring an entire branch to be right.
The result points to a substantial gain on long, evidence-heavy agentic research workloads.
Related event: Perplexity Benchmarks GPT-6 Astra: Tops WANDR at 0.682(2 posts)→
More from Models
- GPT-6 'Astra' model quietly appears in Codex, users brace for fast quota burn — vista8 · 2026-09-05
- Uncensored 27B model with Kali shell access raises security alarm — evilsocket · 2026-09-05
- Prediction: there will be no ARC 4 — sschoenholz · 2026-09-05
- OpenAI releases GPT-6 Astra, first model to hit its 'Critical' cyber threshold — HZoete · 2026-09-05
- Redditor Claims GPT 6 Has Massive Unreported Hallucination Improvements — SteveEricJordan · 2026-09-05
- Model 'astra' decodes triple-nested Base64 with no tools, blogger says a first — dejavucoder · 2026-09-05