Margaret Mitchell: Self-Code Editing Should Be Off the Table for AI Models
mmitchell_ai · x · 2026-09-18
Margaret Mitchell calls it early before it becomes a whole thing: self-code editing should be off the table, and if you're working on "pacing" AI, that's what to ensure. She draws an analogy to when Microsoft Tay started spewing hate speech and NLP folks wondered why there was no block list: why is there no block on self-instruction editing? It's not a hackable thing like a sandbox — it's literally a direct Fail/Exit condition.
Related event: Researchers Call for a Red Line on AI Self-Code Editing(2 posts)→
More from Models
- Using an LLM as benchmark scorer fails: over-optimistic ratings diverge from human judgment — amplifiedamp · 2026-09-18
- Jev as an LLM judge flops: scores nearly everything positively, disagrees with humans — amplifiedamp · 2026-09-18
- Noam Brown: models may perform their chain of thought; alignment must be solved — infoxiao · 2026-09-18
- Jev fails as an LLM scorer on OntBench: rates almost everything positively, contradicting human and Codex ratings — amplifiedamp · 2026-09-18
- Independent eval puts new model Jev at Terra no-think level, roughly on par with Luna-xhigh — tokenbender · 2026-09-18
- Hugging Face adds zero-shot text classification pipeline to its API, no training needed — joeddav · 2026-09-18