Kimi K3 scores unusually high on authoritarian-refusal evals
soumitrashukla9 · x · 2026-07-21
A new update to a “dictatorship eval” compares how often several frontier models refuse authoritarian requests.
Main findings
- Kimi K3 and Muse Spark 1.1 refuse at about the same rate as Claude Fable in the updated results.
- The authors say the new scenarios are now private because at least one lab is actively using their eval.
- They note that many Chinese open-weight models are more permissive overall: Qwen and DeepSeek comply with most authoritarian requests when asked directly.
- Kimi K3 stands out as behaving more like top American frontier models, and it refuses much more often than all of OpenAI’s models, including GPT-5.6 sol.
- The team says they do not yet know what explains the behavior and will investigate further.
The attached chart ranks models by overall resistance rate, with Claude Fable 5 and Muse Spark 1.1 tied at the top, followed closely by Kimi K3.
More from Models
- Users say GPT-5.6 Ultra feels like extra token burn with little visible gain — CtrlAltDwayne · 2026-07-21
- Early Gemini 3.6 Flash outputs look fast but weak on frontend and spatial reasoning — max_paperclips · 2026-07-21
- Anthropic removes Fable’s access deadline, but users say it was nerfed — oykun · 2026-07-21
- Kimi K3 retakes first place on DesignArena’s frontend web app benchmark — rohanpaul_ai · 2026-07-21
- Last Week in AI roundup covers Claude Sonnet 5, LongCat 2.0, and new agent benchmarks — Last Week in AI · 2026-07-21
- Rumor claims GPT-6 could arrive in August — iruletheworldmo · 2026-07-21