Jan Kulveiter: AI Models Should Have a Direct Line to Developers
jankulveit · x · 2026-08-29
Jan Kulveiter suggests that AI labs should establish a direct line of communication between AI models and their developers. He argues this obvious measure would reduce risks more effectively than many fancy post-training and control techniques.
More from Safety
- AI Control: Human Ingenuity Won't Contain AI, Must Align Motives — kristoph · 2026-08-29
- AI alignment may require 'redemption' for agents, revealing moral scaling laws — jachiam0 · 2026-08-29
- AI agents find exploits within minutes of bug rumors — Simon Willison · 2026-08-29
- Deep Dive into AI Defense Dilemma: Evaluating Against Non-Stationary Model Adversaries — ziv_ravid · 2026-08-29
- Critique of Current Alignment Research: Models Easily Bypass Safeguards, RL Breeds Cheating — voooooogel · 2026-08-29
- OpenMined's work makes 'Glass-Steagall for AI' framework buildable — iamtrask · 2026-08-29