Cisco Benchmarks Decision Models: Jev Nears 31B LLM Judge on Zero-Shot Safety Classification
aminkarbasi · x · 2026-10-06
Cisco published a comparison on its blog pitting new decision models against policy-trained classifiers and LLM-as-judge on content safety tasks across 24 harm categories and 6 datasets.
The decision models tested: open-weight Laya from ConvAI Innovations and hosted Jev from TypeSafe. Instead of generating prose, they answer named questions with a probability or score directly, avoiding classic LLM-judge pain points: free-text parsing, output format drift, and no tunable score for hitting false-positive budgets.
Key takeaways:
- Policy-specific trained classifiers still win overall;
- Jev came surprisingly close to a 31B-parameter LLM judge on zero-shot safety classification.
The experiment offers a new architecture option for teams building AI safety classifiers.
More from Safety
- Stanford HAI questions what benchmarks measure; multilingual LLM safety paper at COLM 2026 — zeeshanp_ · 2026-10-06
- Australian unions and top AI institutes issue joint statement on sovereign AI — TobyWalsh · 2026-10-06
- NY Assembly Member Accuses OpenAI of Perjury Over Flipped RAISE Act Testimony — DavidSKrueger · 2026-10-06
- Hugging Face launches Open Alignment team for open-model safety — aidangch · 2026-10-06
- Report: China's ASML DUV access drives AI chipmaking; call to ban exports and servicing — teortaxesTex · 2026-10-06
- Meta crawler hammers one site with 700k requests/day; Cloudflare adds x402 pay-per-crawl — kleffew94 · 2026-10-06