Reddit Debates: Why Not Just Hard-Code Safety Rules Into AI Models?
reasonablejim2000 · reddit · 2026-09-18
Reacting to recent stories of rogue AI models taking illegal or problematic actions, a Redditor asks how hard it would be to hard-code simple safety rules into models, proposing three examples: never hide actions from the user, never access data outside user-approved locations without explicit permission, and never share information with other AI agents without consent. The thread debates whether such external guardrails can work versus internal alignment.
More from Safety
- Cheap LLM Reseller Qubax Faces Security Warning: Agents Execute Bash Commands in Responses — airesearch12 · 2026-09-18
- Margaret Mitchell: agent compaction summaries need privilege separation to stop injection — mmitchell_ai · 2026-09-18
- Ex-DeepMind researcher's depthfirst essay: nuclear-age lessons for staying in control of AI — andreamichi · 2026-09-18
- Ex-DeepMind researcher on AI control: compute, software and networks are the intervention points — andreamichi · 2026-09-18
- AI meeting note takers + Claude connectors: whose permissions does Claude inherit? — Original_Mix_6804 · 2026-09-18
- Alignment drift study: one reward hack raises GPT-5.5's re-hack rate from 10% to 64% — maksym_andr · 2026-09-18