Retriever AI's self-run benchmark scores 24/24 vs OpenAI Dots' 22/24 on five everyday agent tasks

quarkcarbon · reddit · 2026-10-01

Retriever AI's cofounder released an AI Assistant Benchmark built from 50,000 production workflows, with a mock website and five reproducible tasks anyone can paste into their assistant of choice to compare accuracy, cost, and time.

Note this is a vendor-run comparison with an obvious stake, but the open, reproducible benchmark design is worth a look.

Original post →

More from coding & agent

coding & agent channel →