RL-tuned 975B open-weights Inkling gets shorter, steadier dad jokes
shizhediao · x · 2026-07-27
- A user reports RL-training the open-weights 975B MoE model Inkling into PunTune-0.6 using GRPO on Tinker API, so it responds to any topic with a focused original dad joke.
- The key takeaway is not that the model learned new wit, but that training taught it to pick one punchline and stop—improving reliability and efficiency rather than creativity.
- The fine-tuned model reaches strong punchlines in about a quarter of the tokens used by the base model, which the author frames as useful for understanding what training can and cannot unlock in product behavior.
More from Fun
- Reading old books is like reading civilization’s system prompt — Kyrannio · 2026-07-27
- Creator releases an AI-collab song built with two different models — Mpire2025 · 2026-07-27
- Google AI Overview tells a baby sphere to sing, glow, and roll away — Sauers_ · 2026-07-27
- A math student allegedly used an LLM to prove 60,000 theorems in one weekend — burkov · 2026-07-27
- A stats meme turns Simon Wood’s REML smoother into a joke about overfitting — Sauers_ · 2026-07-27
- A post turns Opus 3 into a mythic AI stage actor in a surreal theater riff — repligate · 2026-07-27