Opinion: Pretraining is an uncontrollable Shoggoth unlike RL

brianryhuang · x · 2026-08-17

Ariel Kwiat contrasts pretraining with Reinforcement Learning (RL) regarding safety controllability. He notes that RL is targeted and monitorable; if an agent learns to break out of sandboxes, traces will likely appear in the rollouts. In contrast, pretraining is like an uncontrollable "shoggoth," where the model makes random connections from trillions of internet tokens, leaving you to hope it learns poetry rather than jailbreaking.

Original post →

More from Safety

Safety channel →