Jev beats 7 frontier models on speed and cost ($0.004 vs $1.11), yet all score coin-flip on user preference
PrimaryMagician · reddit · 2026-09-21
Testing a "decision model" (TypeSafe's Jev) against Opus 5, Sonnet 5, Haiku 4.5 and four Codex-routed models on pruning 110 YouTube subscriptions: Jev finished in 9.1s at $0.004 vs 3-7 minutes and $0.03-$1.11 for the rest, with 0.88-0.97 rank agreement with Opus 5 and 3-4x lower run-to-run variance than Sonnet 5.
The key finding: against the author's 66 human keep/drop labels, every model scored AUC 0.43-0.52—a coin flip—while a simple "watched in last six weeks" signal hit 0.69. Models agree with each other but not with the user, whose own blind relabels disagreed 3/13 times. The takeaway build is deliberately boring: model scores plus threshold rules plus human approval via the YouTube API (MIT, github.com/abhibansal60/tidy).
More from Venture
- Intel CEO says company can only meet ~50% of server CPU demand — Beth_Kindig · 2026-09-21
- Vals Builds Private Evaluations Around Real Work, Revenue Up 8x — TansuYegen · 2026-09-21
- Jeff Dean says RL plus new EDA tooling could compress chip design from 2 years to 3 months — ycombinator · 2026-09-21
- Junior ML engineer rejects $400k offer as AI talent market overheats — ingliguori · 2026-09-21
- Dev reports real traction for vibe-coded games on AI-native platform Spawn — majidmanzarpour · 2026-09-21
- If your checkout isn't agent-friendly, agents will just recommend your competitor — jeff_weinstein · 2026-09-21