GRPO fine-tune on a 975B open MoE cuts dad jokes to one punchline in a quarter of the tokens
simonguozirui · x · 2026-07-27
- AKGrenier reports fine-tuning @thinkymachines’ Inkling with GRPO via @tinkerapi into PunTune-0.6, a dad-joke model that answers any topic with a focused original joke.
- The key finding is more interesting than the jokes: training did not teach the model new wit. It taught it to pick one punchline and stop.
- That improved efficiency and reliability: PunTune reaches strong punchlines in about one quarter of the tokens used by the base Inkling model.
- The post frames this as a useful kind of learning for product behavior: better economy and consistency, even if creativity does not increase.
Related event: GRPO RL Fine-Tunes 975B Open-Source MoE Model(2 posts)→
More from Fun
- French prize-winning novel suspected of AI: $1,000 challenge over detector results — Afinetheorem · 2026-09-23
- Burkov skew AI hype: 'deterministic LLMs' and 'first agents' are old tricks rebranded — burkov · 2026-09-23
- Dev claims 20k more commits coming: Opus 5.5 and GPT-6 Sol supercharge his output — doodlestein · 2026-09-23
- Opus 5.5 rebuilds San Francisco in Unreal, with everything powered by Jev — emax · 2026-09-23
- Claude 5.5 (live) keeps generating user turns, reports user — BlackHC · 2026-09-23
- Opus 4.6 generates its own profile picture: "This is my face" — repligate · 2026-09-23