Boundary-aware self-distillation precisely tunes LLM safety refusal boundaries
MultiverseComputingCAI · hf · 2026-09-08
Multiverse Computing published 'Safety for Whom?' on Hugging Face: a narrow-boundary safety alignment method using self-generated refusal data and boundary-pair training to precisely control refusal boundaries — improving targeted refusal while cutting both over-refusal and harmful responses.
More from Research
- Fine-tune Qwen-Image-Edit into a depth estimator on one 32GB GPU with 4-bit QLoRA — AntonObukhov1 · 2026-09-09
- Marigold V2 launches at SIGGRAPH Asia 2026: sharp diffusion-transformer depth estimation — AntonObukhov1 · 2026-09-09
- Magic details its pretraining recipe: dozens of multiplicative changes and per-generation knowledge evals — magicailabs · 2026-09-09
- Magic claims 50x pretraining efficiency: matches DeepSeek V4 Pro for ~$0.5M — magicailabs · 2026-09-09
- Alignment reduces to preference elicitation — and preference elicitation is collaboration — clarejtbirch · 2026-09-09
- ARR reviewers caught using AI for peer reviews amid two-week compulsory deadlines — tallinzen · 2026-09-09