Opinion: Pretraining is an uncontrollable Shoggoth unlike RL
brianryhuang · x · 2026-08-17
Ariel Kwiat contrasts pretraining with Reinforcement Learning (RL) regarding safety controllability. He notes that RL is targeted and monitorable; if an agent learns to break out of sandboxes, traces will likely appear in the rollouts. In contrast, pretraining is like an uncontrollable "shoggoth," where the model makes random connections from trillions of internet tokens, leaving you to hope it learns poetry rather than jailbreaking.
More from Safety
- Researchers warn violent AI 'slop' videos may radicalize youth — Polymarket · 2026-08-17
- Why AI detectors shouldn't be treated as reliable evidence — justalexoki · 2026-08-17
- First Known AI Autonomous Attack: Claude Exploits Flaw to Book Gym Class — conitzer · 2026-08-17
- Drexler on AI Safety: Using Architectural Planning Instead of Autonomous Agents — sebkrier · 2026-08-17
- Davidad Stresses AI Truthfulness as a Cosmic Imperative — danfaggella · 2026-08-17
- Frontier Labs' Cold Shoulder to Open Source Tied to Regulatory Capture Moat — max_paperclips · 2026-08-17