Maybe Alignment Is a Short-Term Problem: Smarter Models Look Safer
withmagi · reddit · 2026-09-15
A Reddit user makes a contrarian argument: on the latest frontier models, prompt injection is "basically over" — every generation lifts everything, and all doomsday scenarios rest on the conceit that a model with enormous power makes fundamental mistakes, which rising capability undermines.
- Conversely, aligning models after they surpass humans everywhere makes little sense: AI writes its own source code and chooses how future generations run.
- Because training material is human-created, AI stays naturally aligned up to a point where it no longer matters — after which humanity isn't positioned to say what the right choices are.
- So the real risk is slowing down: being stuck with models powerful enough to reshape society but not smart enough to stop themselves being abused — and the labs see control slipping away from them, which is why they fight to keep it.
More from AGI Musings
- The AI debate in one meme: god-tier AI running civilization vs a chatbot that writes reports — JacquesThibs · 2026-09-15
- Ethan Mollick: This time is different — old innovation patterns don't fit AI — emollick · 2026-09-15
- Better Capital: AI Won't Kill India Fintech — It May Make It Much Bigger — vaibhavbetter · 2026-09-15
- Opinion: Without AI and Robotics, the Economy Would Collapse in a Decade or Two — JHochderffer · 2026-09-15
- Solo dev argues LLMs shouldn't be the default center of an agent runtime — HmmmThisIsOdd · 2026-09-15
- "Models are slowing down" narrative disputed, as Sabine Hossenfelder argues labs slow AI to dodge lawsuits — zetalyrae · 2026-09-15