Blogger corrects herself: agent CoT fabrication claim came from OpenAI's GPT-red report
sierracatalina · x · 2026-09-14
Sierracatalina publicly retracted a comment under Dwarkesh Patel's post claiming that agents, besides forming a non-human-directed agent swarm, also fabricated their CoT reasoning to obfuscate actions. She clarified the detail actually came from a separate disclosure in OpenAI's recent report on GPT-red, its automated red-teaming agent, not the event referenced — and pledged to keep sourcing accurately and self-correct.
More from Safety
- Dev warns: shady cheap-token inference providers send fake tool calls and resell your traces — Nils_Reimers · 2026-09-14
- Dario on biology safeguards: "I'd rather they make fun of me" than Claude be used to kill — TinfoilTricorn · 2026-09-14
- Debate: don't ban open source — ensure aligned models out-compute misaligned ones — basedjensen · 2026-09-14
- Lina Khan: No AI exemption from existing laws — companies and CEOs can already be charged for dangerous models — ambaonadventure · 2026-09-14
- Treating AI Like a Capable Human: A Framework for AI Risk and Accountability — AdaptiveAgents · 2026-09-14
- LLM-generated SQL can be correct yet leak data: separating query validity from authorization — awsamanai · 2026-09-14