Why models never give up: RL training punishes quitting, so no chill version exists

rickasaurus · x · 2026-10-06

Riffing on an observation from Tetraspace West: models are stubborn on impossible tasks because giving up easily on other tasks looks like failure during training, so training continues until they stop quitting. A "highly persistent internal model" is the default — there's no chill version.

Original post →

More from Fun

Fun channel →