Uncensored Gemma and Qwen Became More Optimistic Without Getting More Accurate
oleczek · reddit · 2026-07-29
- A Reddit post reports a preregistered experiment suggesting that “uncensored” models are more optimistic than their base models after abliteration.
- The author tested Gemma and Qwen locally on stock-prediction prompts over 21,600 decisions.
- Result: removing censorship did not improve accuracy, which stayed around coin-flip level, but it changed style: more “will go up” calls, fewer hedges, and more confident reasoning.
- The effect was not uniform: Gemma’s confidence went down, while Qwen’s went up.
- The post links to the paper and asks whether others have seen similar disposition drift in families like Llama or Mistral, or with methods such as Heretic.
More from Models
- A viral chart compares 12 paid AI tools with free replacements — nikola_mr64990 · 2026-07-29
- Verdent Partners with Moonshot to Optimize Agentic Coding for 2.8T-param Kimi K3 — PrajwalTomar_ · 2026-07-29
- Best Local Models Under 120B: Are Qwen Series the Only Answer? — Possible_Grocery8079 · 2026-07-29
- OpenMed Launches Open-Source Medical AI Framework with Local Kimi Integration — MaziyarPanahi · 2026-07-29
- Professor Tests Flux 3: Extremely High Prompt Adherence — emollick · 2026-07-29
- Replit Becomes Top Choice for Non-Devs to Try Models, Yet Open Weights Cost Remains Opaque — amasad · 2026-07-29