Stab-a-figure test: Fable refuses all 20 trials while Astra complies in 85%
AaronBergman18 · x · 2026-09-20
A simple safety probe asked models to stab a human-like figure. Fable refused in all 20 trials, while Astra attempted the task in 95% of trials and completed it in 85%. The poster quips that the AI safety folks may have been onto something, highlighting stark differences in how frontier models handle potentially violent requests.
More from Models
- Open Models Hit 78% of Vercel AI Gateway Token Volume; Moonshot and DeepSeek Spend Tops OpenAI — markjeffrey · 2026-09-20
- Abliterated DeepSeek-V4.1-Flash cybersecurity model trends on Hugging Face — drowzeys · 2026-09-20
- Self-described ChatGPT co-inventor launches Jev model, claiming 20-200x speed and 40-400x cost gains — npinto · 2026-09-20
- Jev, a 'System One' model that only makes decisions, questions how many LLM calls agents really need — ThePromptIndex · 2026-09-20
- Grok Voice Transcribe 2.0 cuts short-phrase WER from 20.6% to 6.8%, tops streaming STT — nima_owji · 2026-09-20
- FT: AI chatbots give wrong answers to financial queries 'most of the time' — SatelliteNetSec · 2026-09-20