Dev hit with account-termination threat after pasting agent report on his own hack
lucasmeijer · x · 2026-10-06
Developer Lucas Meijer says OpenAI threatened to terminate his account after he pasted in a findings report from the agent that investigated his hacked account — a report describing the attack was apparently flagged as violating content. He calls the experience Kafkaesque: trying to report a security incident nearly got him banned. It's a vivid example of keyword-triggered guardrails failing to distinguish describing malicious behavior from committing it.
More from Models
- CoNLL 2023 paper: instruction-tuned GPT models beat children on Theory of Mind tests — dioscuri · 2026-10-06
- Mistral now testable in Battle and Agent Modes on LMArena — arena · 2026-10-06
- User Claims 'GPT-6' Solved His Favorite CTF Fully Autonomously in About an Hour — SIGKITTEN · 2026-10-06
- Mistral Large 4 burns over 2x the output tokens per task vs GPT-6 sol — haider1 · 2026-10-06
- Mistral insider hails Large 4 release as finally making product plans come together — qtnx_ · 2026-10-06
- Reflection AI Claims Beam Is 3-4x More Inference-Efficient Than GLM 5.2 — ChrSzegedy · 2026-10-06