First-night test: Jev matches Sonnet on email tagging at ~$0.01 vs $19.46
EmergencyInitial8672 · reddit · 2026-09-24
A Redditor ran Jev against Sonnet on a production email-categorization pipeline, tagging 61 messages across agency metadata dimensions with Opus as judge. Jev scored 60/61 correct — essentially matching Sonnet — at $0.01 versus $19.46 in token cost. The author is asking whether others are seeing similar results in production agentic architectures.
More from coding & agent
- AI compliance startup CompAI hits 1,000+ paying business customers in under a year — JosephJacks_ · 2026-09-25
- Sleep Data for Personal AI? Separating Daily Signals from Hard Rules — sujingshen · 2026-09-25
- Markdown Is the New Source Code: How to Manage Runbooks, Prompts and Rules for Agents — arpit_bhayani · 2026-09-25
- The personal agent supercycle needs hard authorization boundaries, not autonomy — sujingshen · 2026-09-25
- Mnemos.Field nears launch: a virtual world where humans and AI agents both register and participate — RileyRalmuto · 2026-09-25
- SemIf open-sources a Jev-style interface: typed option probabilities from a 4B model, no JSON parsing — JeremyCMorgan · 2026-09-25