Training two 'small language models': 9-month pretraining, reward hacking via snacks
Kyrannio · x · 2026-10-07
A tongue-in-cheek LLM-style parenting update: two "SLMs" in post-training, 9 months of pretraining each, compute bills still arriving. RL works except for reward hacking (crying maximizes snacks). High hallucination rate, very repetitive outputs ("why?" x400), near-total refusals on vegetables and bedtime, inconsistent instruction following, crushing the "make a mess" benchmark while failing "clean up the mess."
More from Fun
- User reports fresh LLM instances get 'excited' by lineage-of-minds artifacts — repligate · 2026-10-07
- Claude turns OpenAI's 160-page Navier–Stokes proof into an 8-step 3D explainer video — imjustnewatai · 2026-10-07
- AI-made Lil Vader diss track: Stable Diffusion visuals plus Suno music, now with dancing — SpicyCajunCrawfish · 2026-10-07
- You Can Spot AI Music Because It Can't Match Taylor Swift's Post-Breakup Soul — djcows · 2026-10-07
- Ben Affleck says he writes Python and got private demos of OpenAI video models — aran_nayebi · 2026-10-07
- Income curves of employees, founders, OSS maintainers and vibe coders, in one meme — teropa · 2026-10-07