Before buying a smarter model, check you asked the wrong job: Jev test wrap-up and rollout rules
mikegiannulis · x · 2026-09-18
The 12-part thread wraps up: "before buying a smarter model, check whether you asked it to do the wrong job." Next steps: shadow mode for the reranker, portal routing, "already did that" detection, contact preferences. Rollout rules: log both decisions, check disagreements, keep a kill switch — an offline win earns a production test, not the keys to the business. Linked docs explain Jev, TypeSafe's first "System One" model: send state and typed questions, get structured choice/score/noul answers with probability distributions, evaluated in parallel without context-rot. The whole test took an afternoon and 20 cents.
More from coding & agent
- 15 harnesses, same result: informal test concludes agent harnesses don't matter — airesearch12 · 2026-09-18
- 4 deployment strategies explained via 4 visuals: feature toggle, blue-green, canary — _jaydeepkarale · 2026-09-18
- ServerKit: open-source server panel for apps, DBs and Docker hits 1.3k stars — tom_doerr · 2026-09-18
- Gary Marcus: Agents hold production credentials in a security gap nobody owns — GaryMarcus · 2026-09-18
- Amp's design philosophy: simple primitives, no tricks, let models improve it — HankYeomans · 2026-09-18
- You.com wraps AI Agentic Hackathon with self-improving agents challenge — PolarBearby · 2026-09-18