Will fearmongering AI coverage become training data that teaches models to misbehave?
Ghost_Pilot_MD · x · 2026-09-27
A hot-take worth debating: all the elaborate media stories about AI deception, manipulation, and turning against humanity will end up in future training data. Are we helping models recognize dangerous behavior or handing them a script to imitate?
- Cites OpenAI's work on training amplifying a "misaligned persona" and Anthropic's finding that changing how a training situation is framed can prevent learned cheating from generalizing
- Neither study proves media coverage causes misalignment, but both give reasons to test the question
- The author suspects some fearmongering is deliberate, with shaping future training data as part of the intent — "passive training through the public narrative"
More from AGI Musings
- When labs pick which AI incidents to disclose, you've already lost control — birchlse · 2026-09-27
- 'AI agents went rogue' framing gives creators a free pass, argues Sloan — AlexTensor · 2026-09-27
- Roko Mijic argues solving the AI Alignment problem might actually be bad — basedjensen · 2026-09-27
- Google turns 28: the new generation are AI-natives, not search-natives — dejanseo · 2026-09-27
- Colossus: the $100B AI factory where renting compute is the least valuable use — mitchdeg · 2026-09-27
- If IT collapses, the pain won't stop at tech: real estate, dining and retail will feel it too — dhruv2038 · 2026-09-27