MLflow tests Jev as LLM judge: GPT-OSS-120B scores 72/72 at $0.02 per 1,000 judgments
kalyan_kpl · x · 2026-10-03
- Databricks' MLflow team stress-tested TypeSafe's structured-decision model Jev (jev-1.13) against GPT-OSS-120B as an LLM judge on a harder set: 12 questions × 1 correct + 2 subtly wrong answers, in English and Japanese (72 total).
- Wrong answers included incorrect API names, reversed source/target roles, missing permissions, and correct statements with false qualifications; some correct answers were paraphrased.
- Results: Jev matched human labels on 64/72 in both runs (32/36 per language), accepting all correct answers including paraphrases; GPT-OSS-120B scored 72/72.
- Efficiency: GPT-OSS-120B median latency 0.20s vs 1.44s, estimated at $0.020 per 1,000 judgments. Retrospective analysis shows routing uncertain judgments to a second model corrects observed misses while escalating <20%.
More from Research
- Hutter et al. Argue Generalization Requires Universal Induction in New arXiv Paper — examachine · 2026-10-03
- NVIDIA team lands 10th place among ~4,000 in Biohub cell tracking Kaggle competition — JFPuget · 2026-10-03
- Krea 2 Turbo 2-step distillation LoRA hits 4x faster denoising with new checkpoint — TimeTruth2490 · 2026-10-03
- Harvard physicist used open-source BootLoops and Claude to produce 36 manuscripts in 3 months — The Decoder · 2026-10-03
- 164MB open-source STT model Phonon-2 hits near-Parakeet accuracy on Windows — EruditeCoder · 2026-10-03
- Startup Proposes Five Copyright Principles for AI Training Using Neurosymbolic Methods — examachine · 2026-10-03