Kimi K3 scores unusually high on authoritarian-refusal evals
soumitrashukla9 · x · 2026-07-21
A new update to a “dictatorship eval” compares how often several frontier models refuse authoritarian requests.
Main findings
- Kimi K3 and Muse Spark 1.1 refuse at about the same rate as Claude Fable in the updated results.
- The authors say the new scenarios are now private because at least one lab is actively using their eval.
- They note that many Chinese open-weight models are more permissive overall: Qwen and DeepSeek comply with most authoritarian requests when asked directly.
- Kimi K3 stands out as behaving more like top American frontier models, and it refuses much more often than all of OpenAI’s models, including GPT-5.6 sol.
- The team says they do not yet know what explains the behavior and will investigate further.
The attached chart ranks models by overall resistance rate, with Claude Fable 5 and Muse Spark 1.1 tied at the top, followed closely by Kimi K3.
More from Models
- Developer Building a Unified Leaderboard of All Model Benchmark Scores — airesearch12 · 2026-09-11
- Rumor claims Kimi faked performance by serving Claude; DeepSeek new model surprises in evals — realsohamparekh · 2026-09-11
- GPT-5.6 writes well but is instantly forgettable, user complains — BasedRaddka · 2026-09-11
- Opus Refuses Protein Research Codebase Over 'Safety' Concerns, Dev Considers Rolling His Own — josephdviviano · 2026-09-11
- User Hails Unconfirmed 'DeepSeek 4.1 Flash' as an Inflection Point in LLMs — himanshustwts · 2026-09-11
- Terminal Bench v4: GLM-5.3 Leads at 41.9%, Kimi-K3 Underwhelms at 12.6% — Ok_Warning2146 · 2026-09-11