GRPO fine-tune on a 975B open MoE cuts dad jokes to one punchline in a quarter of the tokens
simonguozirui · x · 2026-07-27
- AKGrenier reports fine-tuning @thinkymachines’ Inkling with GRPO via @tinkerapi into PunTune-0.6, a dad-joke model that answers any topic with a focused original joke.
- The key finding is more interesting than the jokes: training did not teach the model new wit. It taught it to pick one punchline and stop.
- That improved efficiency and reliability: PunTune reaches strong punchlines in about one quarter of the tokens used by the base Inkling model.
- The post frames this as a useful kind of learning for product behavior: better economy and consistency, even if creativity does not increase.
Related event: GRPO RL Fine-Tunes 975B Open-Source MoE Model(2 posts)→
More from Fun
- Musk reposts a 22-year “then vs now” SpaceX meme — elonmusk · 2026-07-27
- Samsung’s DS Division allegedly won’t even supply RAM to its own MX unit — zephyr_z9 · 2026-07-27
- A Denny’s–NVIDIA Venn diagram jokes that both “sell chips” — ZacharyHuang12 · 2026-07-27
- Aethergeist turns an industrial scene into a surreal skull-filled meme — PierceLilholt · 2026-07-27
- In the AI era, “I’ll turn it off” is no longer a serious threat — mallow610 · 2026-07-27
- “Developer after vacation” lands as a simple dev meme — dhruv2038 · 2026-07-27