Dev swaps Gemini 3.1 for Jev as LLM judge: 200x cheaper, 50x faster, same accuracy
indragie · x · 2026-09-25
Developer rbro112 reports one week after moving an eval judging dataset from Gemini 3.1 to Jev by @typesafeai:
- No meaningful change in scoring accuracy
- 200x cheaper ($0.01 → $0.00005 per judge)
- 50x faster (10s → 0.2s median)
- 50% fewer input tokens, 87% fewer output tokens
Not everything is perfect, with details in the original thread. A useful data-point for anyone running LLM-as-judge pipelines.
More from Models
- LightOn launches Ettin model suite: fine-tune a specialized classifier in 42 seconds — IgorCarron · 2026-09-26
- Perceptron Mk1.5 goes live with text, thinking, tools, and audio modes — rohanpaul_ai · 2026-09-26
- AI Detector Flags Enterprise Job Posting as 93% AI, Zero Substance — bushuev_online · 2026-09-26
- Should OpenAI Keep GPT-6 Astra's Successor Internal? The Compute Gap Problem — haider1 · 2026-09-26
- ChatGPT drops text chat limits for free users and upgrades default model — Aiden_Tech_Ai · 2026-09-26
- New Gemini 4 Pro checkpoint spotted in Arena, outranking GPT-6 Astra — cedric_chee · 2026-09-26