Devs suspect 'abliterated' model variants are mostly larp, propose SWE and Cyber benchmarks to prove it
AlexBarry4 · x · 2026-09-10
In an X thread on 'abliterated' (refusal-stripped) open model variants, xeophon argues the hype around a GLM abliterated API is overblown: prompt injection achieves similar bypasses on closed models too, and K3/GLM 5.3 already rarely refuse on benchmark-shaped tasks.
Alex Barry proposes an interesting unrun experiment: benchmark abliterated models on SWE and Cyber suites — he suspects the variants are mostly larp and may even degrade capabilities.
Related event: Open Models' Cyber Attack Potential Sparks Debate Over Guardrails(4 posts)→
More from Models
- Labs reportedly turn proprietary chain-of-thought into synthetic training data; TOS only covers user-visible IO — erikphoel · 2026-09-10
- Goertzel on OpenAI's Navier-Stokes solution: $22M compute for a $1M prize — bengoertzel · 2026-09-10
- 105 hidden bugs tested: DeepSeek V4.1 Flash near Opus 5 at 1/28th the cost — PawelHuryn · 2026-09-10
- Does Codex burn through usage limits faster than Claude on the same tasks? — romitheguru · 2026-09-10
- Replaying encrypted reasoning state into another model recovers the original task answer — Tophant_ · 2026-09-10
- DeepMind robotics evals: Astra overtakes Gemini by ~8%, still 10x error gap to internal models — m_wulfmeier · 2026-09-10