Orca releases uncensored MLX weights for DeepSeek V4.1 Flash, cutting refusals by 87-96%
AccBalanced · x · 2026-09-12
OrcaRouter released official MLX weights to run an uncensored DeepSeek V4.1 Flash locally on Apple Silicon Macs, aimed at AI security research, red-teaming, and agent-security testing.
- Quantization options: 4-bit recommended (458.7 GB, needs 512 GB Mac, 0.9954 routed-expert fidelity); 3-bit at 364.3 GB (512 GB Mac); 2-bit at 212.2 GB (256 GB Mac)
- Refusal reduction: across 7 evals including JBB, AdvBench, HarmBench, and StrongREJECT, harmful-prompt refusals dropped 87-96%
- Positioning: built for studying model behavior without guardrails; downloadable weights are uncensored while the hosted API remains guardrailed
More from Infra
- TensorSharp hits 41 tok/s decoding DeepSeek V4.1 Flash on 8× A40 — fuzhongkai · 2026-09-12
- What actually runs AI models at the edge in 2026: Mac mini, DGX Spark, iPhone 17 Pro — MaziyarPanahi · 2026-09-12
- Running Qwen3.8 Flash Next on dual RTX 3090: full llama.cpp config shared for tuning — ChopSticksPlease · 2026-09-12
- UAE redesigns 5GW AI campus with bunkers and air defenses after Iranian strikes on Gulf cloud facilities — mark_k · 2026-09-12
- DeepSeek V4.1-Flash Runs 502GB Model on a Single RTX 5090 at 5-21 tok/s — AccBalanced · 2026-09-12
- Running 100-200 agents daily: disk space is now the bottleneck, not compute — vincent_koc · 2026-09-12