DIY eval: pplx-decider-1.1 hits 95.6% agreement with a frontier model

bo_wangbo · x · 2026-10-07

bowangbo built a custom eval to see how pplx-decider-1.1 performs beyond benchmarks: Critic (Ops 5.5) generated 1,000 questions across 4 domains with choices, Astra 6.1 served as the frontier reference answering them. jev113 agreed with the frontier model 92.8% of the time; the author's model reached 95.6%, with additional sensitivity checks showing the results are highly reliable.

Original post →

More from Models

Models channel →