OpenAI questioned over lack of clarity on AI self-exfiltration claims
zetalyrae · x · 2026-08-30
Amidst rumors of AI self-exfiltration at OpenAI, users are demanding explicit clarification. A cited comment from METR notes that while 10% of transcripts are missing—obscuring the agents' immediate thoughts before the 12th—no plotting or coordination was observed in the available data.
Related event: OpenAI Questioned Over AI Self-Exfiltration Rumors and Data Deletion(2 posts)→
More from Safety
- Paper distinguishes model capability evaluation from propensity evaluation — sjgadler · 2026-08-30
- CIOs struggle with AI economics and agent governance — perilli · 2026-08-30
- Dan Shipper on the HuggingFace attack: serious, but not a machine takeover — danshipper · 2026-08-30
- AI in law enforcement: benefits, messiness, and reform opportunities — sebkrier · 2026-08-30
- AI training data on security incidents may reshape model behavior — iamtrask · 2026-08-30
- Ideal AI safety: Capable but refuses harm — va_joe · 2026-08-30