DIY eval: pplx-decider-1.1 hits 95.6% agreement with a frontier model
bo_wangbo · x · 2026-10-07
bowangbo built a custom eval to see how pplx-decider-1.1 performs beyond benchmarks: Critic (Ops 5.5) generated 1,000 questions across 4 domains with choices, Astra 6.1 served as the frontier reference answering them. jev113 agreed with the frontier model 92.8% of the time; the author's model reached 95.6%, with additional sensitivity checks showing the results are highly reliable.
More from Models
- Western labs less likely to distill from Claude but more likely from GLM 5.3 — andrew_n_carr · 2026-10-07
- With OpenAI's Luna, cloud beats open weights on price — at least GPUs heat the house — BLUECOW009 · 2026-10-07
- V4.1 Flash review: top of Flash tier but wild hallucination swings, tester wary of V4.1 Pro — teortaxesTex · 2026-10-07
- OpenAI Launches Decisions API in Public Beta, Up to 10x Faster Than GPT-6 Luna — stevenheidel · 2026-10-07
- OpenAI to watermark ChatGPT outputs by default in the EU under AI Act — Ars Technica AI · 2026-10-07
- Mistral trained ML4 on 3,800 Grace Blackwell GPUs in its European datacenter, more clusters coming — rakyll · 2026-10-07