Jev beats GPT Luna at jailbreak detection as a cheap prompt pre-screening filter
mayfer · x · 2026-09-17
Developer mayfer reports that Jev, a small model, is surprisingly good at jailbreak detection, beating gpt luna and serving as a cheap pre-screening filter for prompts. He argues jailbreaks rarely resemble real requests, making a small classifier viable, and expects 100% accuracy on known jailbreak patterns with more tuning.
More from Models
- Burkov predicts looping recurrent 7B transformers will return and get good at coding — burkov · 2026-09-17
- OpenAI's big 'ship week' reportedly postponed, GPT-6 Sol timing now unclear — testingcatalog · 2026-09-17
- Mozilla's 91-page report: open-weight AI now only ~4 months behind the frontier — rohanpaul_ai · 2026-09-17
- Stealth model leak speculated to be Mistral: no output-token billing for reasoning, answers China questions — zainhas · 2026-09-17
- Unreleased Astra-family model reportedly developed a new persona banner during RL training — inductionheads · 2026-09-17
- Researcher Despairs as Gemini Cites 'Emergent Mind' for Made-up AUROC Baselines — anshulkundaje · 2026-09-17