1,565-email benchmark: cheapest Perplexity/Cloudflare decision model is also most accurate
michellechen · x · 2026-10-02
Perplexity and Cloudflare released new decision models, and the author benchmarked them against Jev on the same 1,565 business emails as before: 6,260 calls, 10 categories, 4 models. The cheapest model was also the most accurate at roughly $0.06 per 1,000 emails. Notably, Clef performs far better than its own confidence scores suggest — at 75% confidence it was right 98% of the time. Full benchmark data is public.
More from coding & agent
- A 5-question interview prompt that turns Claude into a scroll-driven page builder — aziz4ai · 2026-10-02
- Single-prompt workflow: Claude Sonnet + GSAP ScrollTrigger builds scroll-driven glass-shatter page — aziz4ai · 2026-10-02
- After agents write customer-specific code, how do you maintain and deploy it? — Embarrassed-Survey61 · 2026-10-02
- Cloudflare launches Web Search API via AI Gateway with Exa, Linkup and Ceramic — michellechen · 2026-10-02
- 20 tasks × 3 repeats = 120 agent runs: the hidden cost of harness comparisons — RelationshipRound711 · 2026-10-02
- Ant's internal Tiger Agent demos Ling-3.1-flash planning workflows across browser, files and terminal — tinkerbellyie · 2026-10-02