Cross-model test says latest Claude refuses a prompt that many frontier models answer
cyb3rops · x · 2026-07-23
A screenshot-based comparison claims that a Hugging Face/OpenAI incident is being over-framed as FUD, and that the latest Claude models refuse to help on a prompt that many other frontier models will answer.
- The author says they ran the same prompt across every model available in OpenRouter and published the results.
- The linked example suggests the latest Claude models refuse, while many other models — including recent OpenAI ones — provide assistance.
- The screenshot shows a benchmark-style dashboard with 342 models, where one model (moonshotai/kimi-latest) is shown as a 10/10 result in the displayed case.
The post is essentially about cross-model safety/refusal behavior on a specific task, and the debate over how much of the incident is model safety versus presentation bias.
More from Models
- Moonshot’s reported Anthropic distillation draws terms-of-service and strategy claims — connoraxiotes · 2026-07-23
- Open-weight Chinese models are still not the same as permissive open source — tamla_htoo · 2026-07-23
- Chinese open-source models may be reinforcing NVIDIA’s inference moat — zephyr_z9 · 2026-07-23
- Kimi generates luxury watch renders with a detailed mechanical movement — cedric_chee · 2026-07-23
- Qwen3.8-Max looks stronger on 3D web frontends than Kimi K3 — cedric_chee · 2026-07-23
- Poster says Opus-5 may beat Fable on cost and performance — adonis_singh · 2026-07-23