Perplexity Open-Sources Agentic Search Eval Environment
andersonbcdefg · x · 2026-07-15
Perplexity has open-sourced WANDR, an eval/RL environment designed to measure agentic search performance.
- The environment is synthesized from production traces to closely mirror real-world distributions rather than acting as a pure toy task.
- It utilizes weak human supervision, making it suitable for evaluating and training stronger models.
- The team stated that they have already used this RL environment internally to train more capable models and plan to scale this approach to other domains and tasks.
- This initiative stems from their product use cases, with the goal of covering a broader distribution of knowledge work.
Related event: Perplexity open-sources internal research benchmark WANDR(9 posts)→
More from coding & agent
- A GLP1R variant may explain stronger Ozempic weight loss, and the team built an agent workflow — julia_kiseleva · 2026-07-21
- A Claude-coded Chrome extension shames you with a private jet when you open YouTube — alex_verem · 2026-07-21
- A curated TTS list for voice agents tracks latency, cancellation, and evals — mahimairaja · 2026-07-21
- Harness engineering is emerging as the execution layer for reliable AI agents — Pavan_Belagatti · 2026-07-21
- DevFest Lisbon keynote will cover Google AI Studio’s latest vibe coding and agentic AI features — gerardsans · 2026-07-21
- Daniel Hanchen’s 2-hour workshop covers open models, reward hacking and RL — danielhanchen · 2026-07-21