Jev loses to Gemini on 1,565-email classification benchmark, but dev still wants it in production
socialwithaayan · x · 2026-09-20
A developer benchmarked TypeSafe's Jev model on 1,565 German and English business emails across 10 categories from industrial suppliers, and Jev lost to Gemini on email classification. Still, the tester remains interested in putting it into production — the interesting result wasn't accuracy but where the mistakes happened.
More from coding & agent
- Are AI agents changing how engineers think, not just how fast they ship? — Relevant-Potential17 · 2026-09-20
- Eight coding agents store sessions eight ways, so one dev built a unified searchable index — Southern_Reference23 · 2026-09-20
- Open-source agent orchestrators may beat model-run orchestration on cost, devs argue — soumitrashukla9 · 2026-09-20
- The degenerate solution problem in agent evals: when writing an end-to-end policy wins — JoshPurtell · 2026-09-20
- Open-source rakazo lets you self-host persistent AI bots with Linux desktops via OpenRouter — dr_cintas · 2026-09-20
- Founder breaks down how to safely wire MCP into a sales outreach agent workflow — namanyayg · 2026-09-20