RL-tuned 975B open-weights Inkling gets shorter, steadier dad jokes
shizhediao · x · 2026-07-27
- A user reports RL-training the open-weights 975B MoE model Inkling into PunTune-0.6 using GRPO on Tinker API, so it responds to any topic with a focused original dad joke.
- The key takeaway is not that the model learned new wit, but that training taught it to pick one punchline and stop—improving reliability and efficiency rather than creativity.
- The fine-tuned model reaches strong punchlines in about a quarter of the tokens used by the base model, which the author frames as useful for understanding what training can and cannot unlock in product behavior.
Related event: GRPO RL Fine-Tunes 975B Open-Source MoE Model(2 posts)→
More from Fun
- Interactive Video Generation in Any Art Style, Pure JS Coded by Claude — ctjlewis · 2026-09-23
- Claude wrote an entire song purely in code, no Suno involved — ctjlewis · 2026-09-23
- Economist worries AI detectors discriminate against aggressively literary French prose — paulnovosad · 2026-09-23
- When scientists can't find trends, they plot the logarithm of all variables — burny_tech · 2026-09-23
- ML author Burkov zings AI hype: 'visionary' claims vs Theranos founder in jail — burkov · 2026-09-23
- Your AI agent ran for 12 hours — but did anyone tell it the brief changed at hour 2? — HaktanSuren · 2026-09-23