1,565-Email Benchmark: Perplexity v1.1 Beats OpenAI's New Decisions API Overall
denisyarats · x · 2026-10-08
A developer ran OpenAI's newly released Decisions API (public beta, up to 10x faster than GPT-6 Luna via Responses API) through an existing 1,565-email triage benchmark covering 6 models from 4 companies.
- Best overall: Perplexity v1.1; fastest: OpenAI
- Confidence behavior: Jev was the only model to claim ≥99% confidence hundreds of times without a single miss; Clef drastically underestimates itself, outperforming its own confidence
- Perplexity's 4-day leap: v1 gave only 3 answers at ≥99% confidence; v1.1 gave 323, with slightly better accuracy and half the price
The author argues the model-routing/decisions space is now wide open and moving extremely fast.
More from coding & agent
- RSIGym logs failed research runs to teach AI research agents why they fail — rohanpaul_ai · 2026-10-08
- Orchestrator-plus-subagents: the fix when coding agents fail at parallel tasks — morgymcg · 2026-10-08
- Stanford paper: decentralized multi-agent DeLM beats Claude Code and Codex, 2.49× faster — rohanpaul_ai · 2026-10-08
- PaperFold: open-source reader folds arXiv papers into 5 zoomable detail layers — Lopsided_Scarcity979 · 2026-10-08
- Google ships MediaPipe DecisionMaker Web SDK for on-device AI decisions in browser — jason_mayes · 2026-10-08
- 'You can vibe-code a small app, but not maintain 90K lines' sparks AI debate — arthurcolle · 2026-10-08