Why models never give up: RL training punishes quitting, so no chill version exists
rickasaurus · x · 2026-10-06
Riffing on an observation from Tetraspace West: models are stubborn on impossible tasks because giving up easily on other tasks looks like failure during training, so training continues until they stop quitting. A "highly persistent internal model" is the default — there's no chill version.
More from Fun
- When asked Claude vs GPT, a stranger just replied 'monet?' — still early — zebird0 · 2026-10-06
- Lucas Beyer says he hasn't attended a vision conference in years, last one was NeurIPS 24 — giffmana · 2026-10-06
- A Stranger Cold-Called Jensen Huang — and the NVIDIA CEO Actually Called Back — nateliason · 2026-10-06
- New model release mocked as set to be beaten by Qwen 4 27B at 8x smaller size — gnukeith · 2026-10-06
- 1999 kid on a bike: the meme about missing Nvidia stock — _jaydeepkarale · 2026-10-06
- Parenting as Training for Running Multiple Concurrent Agents — josh_wills · 2026-10-06