Exa launches ATLAS benchmark: even priciest search agents miss ~1/3 of results
yoimnotkesku · x · 2026-10-09
Search company Exa previewed ATLAS, a new benchmark grading agents on search-intensive real-world workflows, pairing queries grounded in actual search demand with verified golden answers via an automated refresh pipeline to fight benchmark contamination.
Key findings from early runs:
- No agent run under $1 per task achieved a row F1 above 0.5 — cost-efficient exhaustive search is far from solved.
- Even the most expensive search agents miss about one-third of golden results.
- With the harness fixed, Exa claims its own backends define the cost-performance Pareto frontier, with 16% score variance across backends (note: self-interested).
More from coding & agent
- Google unveils universal Gemini agent for work: one prompt box for Q&A, knowledge work and code — sundarpichai · 2026-10-09
- Flipping the AI playbook: deploy agents on the client's own machine, then unplug — curious_vii · 2026-10-09
- Voyager: An Open Agent Harness Built for Creative Work, Not Coding — socialwithaayan · 2026-10-09
- Every's Marketer Has a Slack Agent Open PRs Straight from Figma — every · 2026-10-09
- Devs Ditch 500 .md Files: Rebuild AI Instructions From Scratch Each Model Release — bendee983 · 2026-10-09
- Catalyst raises $30M for autonomous AI trading agent, hiring engineers — jzlegion · 2026-10-09