Replaying 3,300 Real Production Decisions to Test Small Model Jev
mikegiannulis · x · 2026-09-18
The company's product uses AI to help people write books. They replayed 3,300 real production decisions through Jev from @typesafeai, covering tasks like finding sources, routing messages, and catching "I already sent you that."
These are boring tasks, but the ones customers notice immediately when they break — exactly where small model value should be tested.
Related event: TypeSafe Launches Jev, a Decision-Only Model That Never Writes a Word(44 posts)→
More from coding & agent
- Berkeley study: the right agent harness cuts cost of the same result by 71% — MartinGTobias · 2026-09-18
- dotey on the AI code quality debate: black-box verification is replacing code review — dotey · 2026-09-18
- AgentSky launches as an agent marketplace: 40+ coding agents in browser, 44x cost gap between models — Scobleizer · 2026-09-18
- AWS Ships Six Open-Source Skills to Let Coding Agents Deploy Hugging Face Models on SageMaker — AWS ML Blog · 2026-09-18
- Three guys with Claude and Codex subscriptions chained 6 vulns into an advanced attack — xennygrimmato_ · 2026-09-18
- Outerloop 0.2.0 ships persistent agent states and a unified message inbox — mengyer · 2026-09-18