.20 Offline Replay Shows Small Model Jev Beats… · AGI Hunt

$0.20 Offline Replay Shows Small Model Jev Beats Big LLM on Routing and Retrieval Tasks

Mike Giannulis, founder of an AI book-writing company, posted a 12-tweet thread on September 18 fully documenting an offline validation experiment on "replacing a large model with a small one": replaying 3,300 real production decisions through @typesafeai's small model Jev, pulling real inputs from their own database, comparing against past Claude decisions and human-labeled/deterministic ground truth, and pre-registering pass/fail rules before seeing results—fully offline with no online experiments.

Confirmed

Not confirmed

Why it matters

The author's core conclusion: "before buying a smarter model, check whether you're making it do the wrong job." Leaderboards can't reveal issues like "'text, don't call' being filed as other" that only your own data exposes; small models fit "boring tasks where customers immediately notice a screwup." Next step: shadow-mode rollout for the reranker, portal routing, "already done" detection, and contact-preference classification, dual-logging old and new decisions, checking divergences, and keeping a one-click kill switch—an offline validation win only earns a shadow-mode opportunity, not the keys to the whole business.

2026-09-18 ~ 2026-09-18 · 10 related posts

Full story(2 episodes)→

Primary sources