Perplexity Open-Sources Deep Research Benchmark WANDR
AccBalanced · x · 2026-07-15
Perplexity has open-sourced its internal benchmark WANDR, used to build its deep and wide research capabilities.
The quoted content mentions that this benchmark was utilized for the J-space reproduction around GLM 5.2, training reward models, reducing hallucinations via RL, and evaluating models on tasks like cancer prediction.
The core takeaway is that Perplexity has publicly released an internal evaluation benchmark designed to build "deep research" capabilities, demonstrating its application in training and experimental pipelines.
Related event: Perplexity open-sources internal research benchmark WANDR(9 posts)→
More from Research
- Stanford Team Introduces Gigatoken, the World's Fastest Tokenizer — StanfordAILab · 2026-07-22
- Tabul AI launches Metal TreeSHAP to speed up Shapley values on Apple silicon — Scobleizer · 2026-07-22
- Reddit points to OpenAI’s ChatGPT Ads page — EcstaticAsparagus509 · 2026-07-22
- Open-source runtime lets each repo define its own AI code reviewer — ibabufrik · 2026-07-22
- DeepSWE: A New Benchmark for Evaluating AI Coding Agents on Real GitHub Issues — pmz · 2026-07-22
- A Rust space-economy sim runs hundreds of autonomous ships, built with Claude — kalcode · 2026-07-22