Cisco Benchmarks Decision Models: Jev Nears 31B LLM Judge on Zero-Shot Safety Classification

aminkarbasi · x · 2026-10-06

Cisco published a comparison on its blog pitting new decision models against policy-trained classifiers and LLM-as-judge on content safety tasks across 24 harm categories and 6 datasets.

The decision models tested: open-weight Laya from ConvAI Innovations and hosted Jev from TypeSafe. Instead of generating prose, they answer named questions with a probability or score directly, avoiding classic LLM-judge pain points: free-text parsing, output format drift, and no tunable score for hitting false-positive budgets.

Key takeaways:

The experiment offers a new architecture option for teams building AI safety classifiers.

Original post →

More from Safety

Safety channel →