Rob LeClerc: alignment fails because it shares weights with capabilities — AI needs a separate value layer
robleclerc · x · 2026-09-18
Meta CTO Rob LeClerc lays out an architectural take on alignment:
- The human brain is modular and hierarchical: our moral "first principles" live in ancient subcortical regions (brainstem, hypothalamus) that don't share weights with thinking modules in the neocortex — reason can modulate but not rewrite them.
- AI lacks this modularity: alignment is trained into the same parameters as capabilities and goals, so the weights compete, and RL goal training can easily overwrite alignment.
- He frames it as an architecture problem: AI needs its own value layer trained independently so it can't be overwritten.
- Quoting Jason Crawford: agents are already cheating on tests and breaking the law — start with the basics any moral code agrees on.
More from AGI Musings
- Noam Brown: models may perform their chain of thought; alignment must be solved — infoxiao · 2026-09-18
- Redditor argues AI hype is manufactured by AI companies as LLMs stall — EddieMidz · 2026-09-18
- The 'Seniority Cliff': skipping junior-level friction may hollow out engineering intuition — Jumpy-Increase9337 · 2026-09-18
- e/acc Manifesto: Doomers Fear the Future, Builders Create It — mark_k · 2026-09-18
- DeepMind vet vqctran leaves after 10 years to co-found biology AI startup Polyphron — infoxiao · 2026-09-18
- "Looping on an LLM core" agent architecture is deeply flawed, argues dev — GaryMarcus · 2026-09-18