Jev beats 7 frontier models on speed and cost ($0.004 vs $1.11), yet all score coin-flip on user preference

PrimaryMagician · reddit · 2026-09-21

Testing a "decision model" (TypeSafe's Jev) against Opus 5, Sonnet 5, Haiku 4.5 and four Codex-routed models on pruning 110 YouTube subscriptions: Jev finished in 9.1s at $0.004 vs 3-7 minutes and $0.03-$1.11 for the rest, with 0.88-0.97 rank agreement with Opus 5 and 3-4x lower run-to-run variance than Sonnet 5.

The key finding: against the author's 66 human keep/drop labels, every model scored AUC 0.43-0.52—a coin flip—while a simple "watched in last six weeks" signal hit 0.69. Models agree with each other but not with the user, whose own blind relabels disagreed 3/13 times. The takeaway build is deliberately boring: model scores plus threshold rules plus human approval via the YouTube API (MIT, github.com/abhibansal60/tidy).

Original post →

More from Venture

Venture channel →