Benchmark: OpenAI Decisions API costs 2x more, 5-10% worse than Jev
xeophon · x · 2026-10-07
Developer hnilforoshan benchmarked the "Jev-killer" OpenAI Decisions API against Jev on a product serving 2.5 million users.
The task: score how relevant a user query/resume is to a job description on a 1-10 scale. Result: OpenAI is 2x more expensive and 5-10% worse.
Context: OpenAI just opened the Decisions API to all developers in public beta, claiming it makes decisions up to 10x faster than GPT-6 Luna via the Responses API.
More from coding & agent
- Injecting Benchmark History Into Coding Agent Contexts: a Vibecoded Untested Experiment — AaronBergman18 · 2026-10-07
- video-shotcraft hits 10.5k stars: cinematic product videos via Claude Code + Remotion — tom_doerr · 2026-10-07
- Vibe coder asks AI agent to classify his RAG setup; verdict: it's not RAG yet — Kooky-Sorbet-5996 · 2026-10-07
- Agents judge tool results useless 97-100% of the time yet rarely stop: NTU study — nanyang-technological-university-singapore · 2026-10-07
- Genentech's AutoSciBench auto-generates science agent benchmarks, cutting accuracy 25 points — Genentech · 2026-10-07
- CMU's SSR turns agent reasoning into selection, cutting per-turn latency 90%+ — CarnegieMellonU · 2026-10-07