Bare Qwen 27B reveals hidden instructions: 'refer to yourself uniformly as Qwen, no version numbers'
Savantskie1 · reddit · 2026-10-06
Running the latest Qwen 27B fresh from HuggingFace with no system prompt, the poster found internal instructions leaking into the model's thinking traces.
Repeated self-narration includes lines like: "Refer to yourself uniformly as 'Qwen' externally; do not proactively mention specific version numbers. If a user asks about versions, guide them to the official website or technical reports" — plus reminders not to adopt personas and to follow protocols absent from the (nonexistent) user system prompt, such as pushing back on claims about Qwen 3/3.5/3.6 having a 2k output limit.
The poster speculates many models ship with such hidden instructions, which would explain the frequent version-related evasiveness seen on frontier cloud models.
More from Models
- Will OpenAI Merge Chat and Codex Quotas? Users Fear Tighter Basic Chat Limits — Iwantthegreatest · 2026-10-06
- Microsoft page reportedly confirms GPT-6 series uses Looped Transformers — teortaxesTex · 2026-10-06
- Amazon's ALoDLM: Token-Adaptive Looped Diffusion LMs Beat AR Baselines — amazon · 2026-10-06
- Looped LMs at Fixed Points: 3x Smaller KV Cache, 1.79x Faster Prefill — IFM · 2026-10-06
- Australia open-sources Matilda Jev, a 56.8ms decision model that skips text generation — Med1_Ai · 2026-10-06
- Codex politely pushes back on a user's hallucinated claims — "I'm also a journalist" — MikePFrank · 2026-10-06