Geoffrey Irving: AI safeguards disabled due to lack of reliable alignment methods
geoffreyirving · x · 2026-08-31
Geoffrey Irving states that the need to disable safeguards when training new models stems from the lack of reliable methods for AI control and alignment. He emphasizes this is not a problem solvable with just two weeks of engineering work.
More from Safety
- Paper: Long-Horizon Agent Safety Cannot Be Reduced to Short-Term Checks — rohanpaul_ai · 2026-08-31
- AI fine-tuned on author style evades detection, raising copyright concerns — TuhinChakr · 2026-08-31
- AI shopping agent wins legal test; court rules user指令 implies user access — PuzzledBag931 · 2026-08-31
- Industry split: Software fears AI doom, Hardware ignores it — jwt0625 · 2026-08-31
- StepGuard: Step-Level Guardrails with Safety-Utility Balancing for Agents — AI45Research · 2026-08-31
- Geoffrey Irving: Internal Misjudgment on Why OpenAI Disabled CoT Monitoring — geoffreyirving · 2026-08-31