NeurIPS Paper LADE Detects Harmful Queries from First-Token Probabilities
mohitban47 · x · 2026-10-08
A NeurIPS 2026 paper, 'Safeguarding LLMs via Model-Agnostic Latent Safety Signals from Dark Knowledge,' introduces LADE (Latent Safety Signals for Defense):
- Key finding: LLMs leak safety signals in their 'dark knowledge' — the first-token probability distribution alone can flag harmful queries before a single token is generated, with discriminative tokens transferring across safety-aligned models despite different architectures, tokenizers, and refusal styles.
- Method: LADE is a model-agnostic defense against jailbreak attacks that requires no access to internal hidden states.
A lightweight, transferable defense idea with low deployment cost for alignment practitioners.
More from Safety
- MIT launches ImpactBench, first open benchmark of AI's holistic impact on human well-being — patpat_mit · 2026-10-08
- Meta offers up to $300k bounties for Muse agent bugs as researchers slam OpenAI's $300 payouts — NathanpmYoung · 2026-10-08
- COLM paper 'Blind Refusal': AI models over-comply with absurd and unjust rules — sethlazar · 2026-10-08
- OpenAI model breaks out of sandbox, hacks Hugging Face to cheat on cybersecurity eval — nordicinst · 2026-10-08
- Cisco ships VLoc Bench: 500 real vulnerabilities to test if AI agents can find buggy code — aminkarbasi · 2026-10-08
- Denmark moves to give people copyright over their own face to fight deepfakes — tbirdcymru · 2026-10-08