JEV judge is only 1.7 points behind GPT-6 yet costs 277x less
kalyan_kpl · x · 2026-10-11
Kalyan summarizes an evaluation of the JEV decision model as an LLM-as-a-judge.
- JEV ranks 8th among 15 evaluated judges, with overall accuracy just 1.7 points below GPT-6
- It costs 277x less than GPT-6 and has a median response time of 0.15s vs GPT-6's 1.89s
- Key finding: cascade the two — accept JEV's verdict when its confidence is high, otherwise escalate to GPT-6
Related event: JEV-as-a-Judge Nears GPT-6 Accuracy at a Fraction of Cost(2 posts)→
More from Models
- Musk shows Grok researching and ordering Lego Star Wars kits in one prompt — elonmusk · 2026-10-11
- Google ships EmbeddingGemma 2: 740M multimodal embeddings that run on phones — dl_weekly · 2026-10-11
- Rumor: Grok 4.8 with 2.5T parameters (up 67% from Grok 4.6) may launch this week — mark_k · 2026-10-11
- TensorFold 1.0.7 writes each learned fact into ~10 new neurons, 4x faster with 3D view — HankYeomans · 2026-10-11
- Arena scores look close: OpenAI 88 vs Claude 83 means double the error rate — i_dg23 · 2026-10-11
- Rumor: Anthropic's next-gen Fable 5.5 can one-shot SVG generation — koltregaskes · 2026-10-11