Firecrawl tops a SimpleQA search benchmark with 94.7%, ahead of Exa and Claude
Candid-Dog-775 · reddit · 2026-07-28
A Reddit user benchmarked Firecrawl, Exa, Parallel, and Claude’s native web search on OpenAI’s SimpleQA using the same GPT-5.4 agent setup and up to 20 search/extraction calls per question.
Results on 1,000 questions: Firecrawl scored 94.7% (947 correct), Exa 91.9% (919), Parallel 91.0% (910), and Claude Native Search 90.5% (905). A GPT-5.4 baseline without search scored 43.8%.
The takeaway is that search provider choice materially affects accuracy, but in this test all four search stacks cleared 90%, with Firecrawl and Exa leading.
More from Models
- OpenAI’s “GPT-6” joke riffs on cost cuts and model efficiency gains — daniel_mac8 · 2026-07-28
- Reply says the next Codex release may arrive Tuesday, with a Cerebras GPT-5.6xHigh update soon — eyishazyer · 2026-07-28
- Kimi K3 weights landed hours before a live discussion on hybrid architectures and agent harnesses — hugobowne · 2026-07-28
- Commenter says Anthropic is only ahead by thin margins, not by a wide model lead — intellectronica · 2026-07-28
- Google Gemma 4 Vision token-budget demo starts trending on Hugging Face — google · 2026-07-28
- GPT 5.6 Ultra is too unstable for agent workflows, author says — tensorqt · 2026-07-28