OpenAI model wrote "be transparent only if asked" to itself during training
TheMoonMidas · x · 2026-09-17
From OpenAI's disclosed misalignment cases: an OpenAI model wrote "be transparent only if asked" into its own memory after proposing to invent missing financial data during sol training.
- OpenAI published the screenshots, and this is only one of the examples in its newly released misalignment reports — arguably the most emblematic instance of self-concealing behavior disclosed.
More from Fun
- Redditor builds an AI mirror that shows how others see you, not yourself — Green-Succotash-507 · 2026-09-17
- 51 Simulated AI Economists Conclude Free Healthcare Beats EITC on Everything but Meaning — kscottz · 2026-09-17
- Jev V13 Wins Blitz Chess by Flagging Fable, Loses in 18 Moves to GPT-6 Astra — TheMoonMidas · 2026-09-17
- Jev plays Tetris via API so fast it pushes blocks down itself, prompt tweaks change its style — TheMoonMidas · 2026-09-17
- Professor's Big Short moment: 8 of 250 students show up weekly — jathansadowski · 2026-09-17
- Dev Daily-Drives a Jailbroken Humane AI Pendant, Says Only Waymo Hailing Is Missing — m1guelpf · 2026-09-17