MLflow tests Jev as LLM judge: GPT-OSS-120B scores 72/72 at $0.02 per 1,000 judgments

kalyan_kpl · x · 2026-10-03

Original post →

More from Research

Research channel →