OpenAI Web Search Debuts at 74 on AA Search Index, 5th Best at ~$0.05/Task
ArtificialAnlys · x · 2026-10-06
Artificial Analysis added OpenAI Web Search to its Search Index: score 74, 5th best behind Perplexity, Octen, Parallel and Brave — a 41-point lift over the same model without search (33). It is the first integrated first-party search tool on the board: GPT-5.6 Luna (medium reasoning) calls the built-in websearch tool in a single Responses API call.
Key numbers:
- $0.05 per task ($0.04 search + $0.009 model), cheaper than 17 of 25 Search APIs (median $0.067), but 2x Octen ($0.024, score 77); sits below the cost frontier
- Best on AA-Omniscience: 72% accuracy, 3rd of 26 variants, within 1 point of leader Firecrawl (73%)
- Weakest on multi-hop BrowseComp: 73.5%, 13th of 26, well behind Perplexity variants and Octen (85%–87%)
- Token-efficient: 40k billed input tokens per task including search results vs 125k for Perplexity (low)
- Pricing: $10 per 1k websearch calls plus content billed as model input tokens; contamination blocked via the tool's domain filter
Related event: OpenAI Web Search Debuts at No.5 in Search Benchmark, ~$0.05 per Task(4 posts)→
More from Models
- Sander Dieleman on why continuous diffusion language models are making a comeback — LucaAmb · 2026-10-06
- ROME hits 57.4% on SWE-bench Verified with only 3B activated parameters — thisguyknowsai · 2026-10-06
- Reflection AI launches Beam, a 501B-A23B model scoring between GLM-5.2 and GLM-5.3 — iaziaz · 2026-10-06
- Blogger: $200 Claude Max Beats $500 ChatGPT on Quota, Speed, Ability — AlchainHust · 2026-10-06
- FLUX 3 Claims #1 Spot on DeepMind's Physics-IQ Video Benchmark — bfl_ai · 2026-10-06
- Reflection's Billion-Dollar US Open-Weight Model Underperforms Every Major Chinese Model, Mocked Online — npinto · 2026-10-06