Benchmarking OpenAI's Decisions API against Jev, Clef and Polymarket — models copy market prices instantly
Physical_Pepper6294 · reddit · 2026-10-08
An open-source experiment called Polyseer pits OpenAI's new Decisions API against TypeSafe's Jev and Cloudflare's open-weights Clef on predicting live Polymarket/Kalshi markets, stripping all odds-quoting sources from the evidence. Key finding: models instantly copy the market when prices appear in context (0.63 → 0.71 when told "market prices 71%"), so 'AI vs market' comparisons are meaningless without filtering. The three brains reason very differently — Decisions is decisive, Jev barely moves, and Clef was swung from 8% to 63% by an SEO spam page. Per-article replays show exactly which headline moved which model; 27 decisions per question run in 2 seconds. Built with Next.js 16, Cloudflare Workers AI and Valyu search.
Related event: OpenAI's Decisions API Falls Short of TypeSafe's Jev in Benchmarks(4 posts)→
More from coding & agent
- Cloud agent runtime plus Tailscale: the "local hands" setup devs prefer — blelbach · 2026-10-08
- Rocket League-style game built with PlayCanvas and Claude runs inside Snapchat as a download-free Lens — willeastcott · 2026-10-08
- PyTorch's ezyang: Claude knows variational types and implements them for you — ezyang · 2026-10-08
- repo2graph turns codebases into graph context for coding agents — and shows where grep still wins — JeremyCMorgan · 2026-10-08
- New multi-agent swarm writeup: smarter models benefit more from collaborative work — repligate · 2026-10-08
- VibeBuddy is a $59 desk gadget that watches your coding agents and speaks up when they need you — juntao · 2026-10-08