GPT-6 Astra tops Perplexity's WANDR benchmark at 0.682, 13.5% above Fable 5.1 at 6.1% lower cost

rohanpaul_ai · x · 2026-09-05

Perplexity published its evaluation of GPT-6 Astra on the WANDR benchmark, scoring 0.682 at $11.98 per task — the highest of any model tested.

Key numbers:

WANDR is unusual in testing wide-and-deep research: finding large sets of qualifying entities and backing every fact with checkable evidence. It contains 500 public tasks requiring 170,495 source-backed records, penalizing incomplete research even when found facts are correct, with strict hard scores requiring an entire branch to be right.

The result points to a substantial gain on long, evidence-heavy agentic research workloads.

Related event: Perplexity Benchmarks GPT-6 Astra: Tops WANDR at 0.682(2 posts)→

Original post →

More from Models

Models channel →