Before buying a smarter model, check you asked the wrong job: Jev test wrap-up and rollout rules

mikegiannulis · x · 2026-09-18

The 12-part thread wraps up: "before buying a smarter model, check whether you asked it to do the wrong job." Next steps: shadow mode for the reranker, portal routing, "already did that" detection, contact preferences. Rollout rules: log both decisions, check disagreements, keep a kill switch — an offline win earns a production test, not the keys to the business. Linked docs explain Jev, TypeSafe's first "System One" model: send state and typed questions, get structured choice/score/noul answers with probability distributions, evaluated in parallel without context-rot. The whole test took an afternoon and 20 cents.

Related event: Small model beats LLM on reranking and routing; real win is latency, not the 20-cent bill(8 posts)→

Original post →

More from coding & agent

coding & agent channel →