OpenAI and Cloudflare launch decision APIs — tested at 355 decisions, TypeSafe's Jev still decides more
PawelHuryn · x · 2026-10-08
OpenAI and Cloudflare launched decision APIs — no text output, just probabilities over options. The author benchmarked them against TypeSafe's Jev across 355 decisions each: classifying 50 business documents into 6 types, routing 24 support messages by tricky house rules, and moderating audience names in a live Q&A app.
Findings:
- Documents and explicit-rule routing: near-perfect across the board (50/50, 24/24).
- The real gap was autonomy: with a ≥0.9 confidence-gap threshold for auto-deciding, of 43 names Jev decided 31, Cloudflare Clef 22, OpenAI 8, Clef-flash 4.
- No model let a bad name through, but OpenAI and Clef-flash would have forced retyping names like "Bo" and "Michael Chen".
- Gotcha: Clef reports the gap squared on yes/no questions — reading it as-is means deciding 3 names, not 22. "0.9" isn't the same 0.9 everywhere.
The author stayed with Jev for best cost-efficiency and fewest human escalations.
Related event: OpenAI's Decisions API Falls Short of Jev in Benchmarks(4 posts)→
More from coding & agent
- LangChain ships Managed Deep Agents v0.9 with agent-created schedules and per-run config — LangChain · 2026-10-08
- Vercel's AI SDK hits 30 million weekly downloads, up 6x in under a year — lgrammel · 2026-10-08
- GitHub HydraFusion adds local model routing as Microsoft ships MAI Code 1.1 at 3-bit, 256K — BenBajarin · 2026-10-08
- Devin Mobile hands-on: cloud agents running on Linux, macOS, or Windows from anywhere — DevinAI · 2026-10-08
- GodTerm: open-source tool pools multiple Claude Code and Grok accounts with limit rollover — Daniel_Farinax · 2026-10-08
- Jeffrey Emanuel's "say no to process" agent skill kills Codex ceremony output — used hundreds of times a day — doodlestein · 2026-10-08