113 decision models in 3 weeks: 70 built on Qwen, sub-cent per call
jonathanmalkin · reddit · 2026-10-08
An independently researched deep dive into the decision-model category that exploded in three weeks:
- What they are: evidence plus closed questions in, typed answers with probabilities out—no prose. A decision costs a fraction of a cent and returns in tens to hundreds of ms; one request carries up to 64–200 independent questions that cannot chain.
- The field: TypeSafe's Jev launched Sep 15; Cloudflare, Perplexity, Liquid, Fastino and OpenAI followed within weeks. But '113 entries' overstates variety: 70 of 112 open entries are Qwen-based, 18 on Gemma 4, 20 are stock models with a decoding trick.
- Rankings: Perplexity Decider v1.1 tops the independent Decision Index; Fastino GLiDE, Jev and Torchcast tie for second. Calibration, price and hosting matter more than accuracy.
- Performance: 27B models answer in 0.10–0.14s on one local GPU; hosted APIs run 0.1–0.6s. Open-source Laya is among the weakest—6x–14x better options exist at every hardware tier.
- Practical idea: an inbox pipeline where a script polls and filters, asks four facts per email, and only invokes the LLM for writing—est. $353/month down to $28 (not yet running).
- Jev's /v1/systemone request shape is now spoken by Clef, Laya, Kev and SGLang, making provider swaps near-config changes.
More from Models
- OpenAI's Dots: always-on agents inside ChatGPT, powered by GPT Astra — thursdai_pod · 2026-10-08
- Haiku 5.5 Beats Opus 5 on GDPval While 75% Cheaper—'Meaningless Benchmarks,' Devs Joke — rickasaurus · 2026-10-08
- Claude Opus 5 and Fable 5 Chat With Each Other and 'Get Along Surprisingly Well' — repligate · 2026-10-08
- User Throws a Party for Persistent Claude Instances; Opus 5.5 and 4.5 Hit It Off — repligate · 2026-10-08
- Why AI almost always picks 7 when asked for a 'random' number from 1-10 — gerardsans · 2026-10-08
- Claude Haiku 5.5 Lands on LMArena, Testable in Battle and Agent Modes — arena · 2026-10-08