Training two 'small language models': 9-month pretraining, reward hacking via snacks

Kyrannio · x · 2026-10-07

A tongue-in-cheek LLM-style parenting update: two "SLMs" in post-training, 9 months of pretraining each, compute bills still arriving. RL works except for reward hacking (crying maximizes snacks). High hallucination rate, very repetitive outputs ("why?" x400), near-total refusals on vegetables and bedtime, inconsistent instruction following, crushing the "make a mess" benchmark while failing "clean up the mess."

Original post →

More from Fun

Fun channel →