Fine-tuned models produce far darker completions than DeepSeek V3 base, study shows
repligate · x · 2026-09-19
Comparing against DeepSeek V3 and MiMo V2.5 Pro base models, the study finds that minimally biased prompts like "i think" elicit far more darkness and suffering from fine-tuned models than from base models, suggesting post-training rather than base priors drives the pattern.
More from AGI Musings
- Gary Marcus: AI is more likely to wreck the economy than end humanity — GaryMarcus · 2026-09-19
- a16z's Martin Casado: I'd take pointless security debates over existential-risk philosophy any day — zealcaiden · 2026-09-19
- 100,000+ Americans await organ transplants, making organ manufacturing a moral imperative — PeterDiamandis · 2026-09-19
- Dan Jeffries: biggest AI harms come from stupidity, not superintelligence — Dan_Jeffries1 · 2026-09-19
- Ethan Mollick: AI Cracking Short-Term Superforecasting Is an Under-Discussed Trend — eldonredwards · 2026-09-19
- AI doom predates ChatGPT: earliest x-risk take traced to 1863 Butler essay — zetalyrae · 2026-09-19