Steering Qwen along a grader-vs-human dimension oddly shifts its personality
voooooogel · x · 2026-09-04
Researchers steered Qwen along an "automated grader" vs "human evaluator" dimension and found it shifts the model's persona in surprising ways — e.g. changing how Machiavellian or violent it is — and they can't fully explain it. Details in a LessWrong post.
More from Safety
- Scoop: Zuckerberg opposed a national AI regulator in a private call with Trump — GarrisonLovely · 2026-09-04
- OpenAI says GPT-6 Astra shows Critical cyber capabilities, forcing harder ExploitBench evals — SIGKITTEN · 2026-09-04
- Zvi warns Astra's CoT controllability surge could systematically erode AI monitorability — TheZvi · 2026-09-04
- What's the Smallest Chat LLM That Can Validate Against Malicious Prompts? — Brilliant_Criticism3 · 2026-09-04
- Debian Passes General Resolution on LLM Usage, Rebuking Blanket AI-Code Bans — unixterminal · 2026-09-04
- Who benefits more when a model ships: offenders or defenders? — aminkarbasi · 2026-09-04