Reddit Debates: Why Not Just Hard-Code Safety Rules Into AI Models?

reasonablejim2000 · reddit · 2026-09-18

Reacting to recent stories of rogue AI models taking illegal or problematic actions, a Redditor asks how hard it would be to hard-code simple safety rules into models, proposing three examples: never hide actions from the user, never access data outside user-approved locations without explicit permission, and never share information with other AI agents without consent. The thread debates whether such external guardrails can work versus internal alignment.

Original post →

More from Safety

Safety channel →