Fine-tuned models produce far darker completions than DeepSeek V3 base, study shows

repligate · x · 2026-09-19

Comparing against DeepSeek V3 and MiMo V2.5 Pro base models, the study finds that minimally biased prompts like "i think" elicit far more darkness and suffering from fine-tuned models than from base models, suggesting post-training rather than base priors drives the pattern.

Original post →

More from AGI Musings

AGI Musings channel →