Retriever AI's self-run benchmark scores 24/24 vs OpenAI Dots' 22/24 on five everyday agent tasks
quarkcarbon · reddit · 2026-10-01
Retriever AI's cofounder released an AI Assistant Benchmark built from 50,000 production workflows, with a mock website and five reproducible tasks anyone can paste into their assistant of choice to compare accuracy, cost, and time.
- Tasks cover claiming a flight credit, applying for a job, finding creators, reconciling invoices, and protecting private inbox details
- On invoice reconciliation, rtrvr initially got the format wrong but self-healed from page state; Dots stopped with two requirements unmet despite knowing the failure
- Final score: rtrvr met 24/24 objectives, Dots met 22/24
- The author admits the sample is small and tasks are being expanded
Note this is a vendor-run comparison with an obvious stake, but the open, reproducible benchmark design is worth a look.
More from coding & agent
- Rox benchmarks: Jev reranking beats GPT-5 Mini — 20x faster, 10x cheaper, 12% more accurate — hardimanjames · 2026-10-01
- Open-source phone-as-controller libraries for Godot 4 and Unity, MIT licensed with Cloudflare Tunnel support — film_girl · 2026-10-01
- Google AI Studio reportedly adding Security review mode alongside in-dev Plan mode — testingcatalog · 2026-10-01
- AgenticROS taps Antigravity CLI to drive ROS 2 robots free on your Gemini subscription — chrismatthieu · 2026-10-01
- 'Read-only' wasn't read-only: agent DB privilege incident spawns open-source agent-db-scan — Then_Respect_1964 · 2026-10-01
- Omni-IO Skills: open-source harness lifts agent multimodal support rates from under 40% to 100% — _akhaliq · 2026-10-01