Testing whether "abliterated" models are just larp: run them on SWE and cyber benchmarks
xeophon · x · 2026-09-10
- xeophon proposes an experiment: run "abliterated" (safety-layer-removed) models on SWE and cyber benchmarks. He suspects they are mostly larp and/or that abliteration actually degrades capabilities.
- The thread ends with an agreement to check back in 5-6 months, leaving it as a testable prediction.
More from Models
- DeepSeek V4.1 Hailed as the 'First Gamer Model' After Blowing Away an FPS-Generation Test — teortaxesTex · 2026-09-10
- Bug Hunt Bench grades frontier models on 105 real bugs; DeepSeek-V4.1-Flash lands 24/105 for $1.80 — PawelHuryn · 2026-09-10
- RSI is here, just disaggregated: DeepSeek using LLMs to design algorithms — teortaxesTex · 2026-09-10
- DeepSeek V4.1 Flash nears Opus-class quality in motion video tests — mesmerlord · 2026-09-10
- Google's last Pro model shipped all the way back in February — ChrisGPT · 2026-09-10
- Anthropic models showed extreme bias toward prior beliefs in hacking incidents — asusarla · 2026-09-10