Openbenchmarks ranks web search APIs for coding agents; Tiny_Fish takes #1 on cost with 46% fewer tokens

Scobleizer · x · 2026-09-10

Openbenchmarks released the "Web Search for Coding Agents" benchmark (2026): 100 realistic product tickets where the agent must patch inherited code by retrieving the vendor's official docs—key implementation details are never revealed in the ticket, patches must compile, and cited URLs must genuinely appear in that run's search results, blocking memorized answers.

Key insight: coding is the #1 AI agent use case, but most token budget goes to searching for API docs, changelogs, and best practices—not writing code. Search is where cost and quality collide: cutting tokens hurts accuracy, protecting accuracy balloons the bill.

TinyFish climbed from 7th to 3rd on search-and-fetch, ranking #1 on token efficiency and cost per task, overtaking Perplexity and Parallel Web Systems and closing in on Exa. After a revamp two weeks ago, task completion rose from 69% to 79%—just 4 points behind leader Exa Deep—while using 46% fewer tokens and costing 60% less per task.

Original post →

More from coding & agent

coding & agent channel →