OpenRouter Launches Web Search Benchmarks for Models and Configurations
AravSrinivas · x · 2026-08-14
OpenRouter announced Web Search Benchmarks, providing rankings of search tools across different models and configurations to help decide how to ground agents. Benchmarks include τ²-Bench Airline, GPQA Diamond, BrowseComp, DeepSearchQA, and HLE. Each score links to configuration, costs, and telemetry. For example, in BrowseComp, Perplexity with Claude Opus 5 (high) leads with 89.0% quality; in DeepSearchQA, Perplexity with Claude Opus 5 (high) leads with 76.5%.
More from Research
- AI for Science Symposium on 8/27 to feature talks on AI accelerating discovery — _shreya_s · 2026-08-15
- High School Student Uses AI to Identify 1.5 Million Potential Variable Stars in NASA Data — PeterDiamandis · 2026-08-15
- 27B agent Faraday beats Claude Opus 4.8 and GPT-5.5 on research replication via new Replica method — omarsar0 · 2026-08-15
- AI Reprocessing of 80K PDB Structures Reveals 60K+ Protein Conformational Ensembles — anshulkundaje · 2026-08-15
- Neuro-symbolic world models gain support: top ARC-AGI-3 harnesses use this approach — GaryMarcus · 2026-08-15
- CMU PhD on adaptive AI agents: from memory & skills to context adaptation, what's the challenge? — Diyi_Yang · 2026-08-15