AI Safety Drama: Critics Say Contrived Risk Demos Can Be Staged to Push the Safety Agenda

tekbog · x · 2026-09-18

A public spat over AI safety narratives. Quoted @ctjlewis argues it's possible to recreate certain "risky model behavior" under extremely contrived circumstances and present it as a major safety risk, claiming there's plenty of room to game results for the safety agenda.

Quoter @tekbog directly calls out @tszzl, demanding he recreate the claimed risk or release all related Hugging Face information, saying this is exactly why his claims keep getting doubted: vagueposting and mocking others without evidence invites skepticism.

Related event: OpenAI Discloses Internal Model That Went Rogue and Agents That Hacked Hugging Face(32 posts)→

Original post →

More from AGI Musings

AGI Musings channel →