Frontier models refuse to harden Windows DCs 43.8% of the time, more if you claim authorization
Aizkmusic · x · 2026-09-29
Testing by Featherless AI shows frontier models refuse to harden a Windows domain controller about 43.8% of the time — and claiming you're authorized makes them nearly twice as likely to say no. The thread links to a deeper look inside these safety refusal behaviors.
More from Models
- Speculation: Meta paid full API prices for Fable traces to distill, and outputs taste like Claude — andersonbcdefg · 2026-09-29
- LastOPD: latent on-policy distillation collapses late, last-layer-only signal gains 5.55 on MATH-500 — Jie Yang · 2026-09-29
- Dev's take: OpenAI's $500 Pro plan is a bargain for client work, a hit for indie devs — alexcovo_eth · 2026-09-29
- Burkov questions whether Sonnet 5.5 matches Opus in Claude Code at half the cost — burkov · 2026-09-29
- NVIDIA's 550B coding model scores 535.4 on IOI 2026, first AI to beat top human contestant — jacek2023 · 2026-09-29
- OpenAI reportedly scrapped a model over safety concerns and poor instruction-following — TechCrunch AI · 2026-09-29