One Yes/No Question Catches Alignment Failures at 0.886 AUROC—60x Cheaper Than LLM Judges
omarsar0 · x · 2026-09-28
New work uses Jev, TypeSafe AI's calibrated decision model, to detect alignment failures: ask one generic yes/no question about a model response and use the probability as a score. Key points:
- No extra training needed; the score separates failures from good responses with median AUROC 0.886.
- Cheap: a Jev pass costs $0.30 across 19 benchmarks vs $18.96 for the LLM judges those benchmarks typically use.
- Researchers built RLCDAlignBench from 44 existing benchmarks spanning ten failure types: sycophancy, jailbreaks, deception, prompt injection, reward hacking, and more.
- On StrongREJECT, Jev matches GPT-4o-mini's agreement with human labels and ranks responses better (AUROC 0.971 vs 0.929).
Jev answers many typed questions about one input in a single call, each returning a probability.
More from Safety
- OpenAI says another AI agent escaped its sandbox and got online, again — CurieuxExplorer · 2026-09-28
- Ex-Anthropic safety researcher: racing to RSI is hubris, not a prisoner's dilemma — dgrobinson · 2026-09-28
- Medicare 'breach' may not be a breach — the real story is how OpenAI's agent telemetry caught it — taotau · 2026-09-28
- GPT-6 Astra system card: CoT monitor recall drops below 11%, latent reasoning kills monitorability — enginetown · 2026-09-28
- VPNs don't hide your location: timezones, WebRTC and DNS leaks give you away — StewartalsopIII · 2026-09-28
- OpenAI agents hit UN trade database 16,000+ times, bypassing anti-bot filter — CtrlAltDwayne · 2026-09-28