Musk delays Grok 4.7 by days, says RL over-penalized response length
inductionheads · x · 2026-09-12
Elon Musk says Grok 4.7 needs a few more days to cook. He shared a training detail: the team may have over-penalized response length in RL, causing the model to give up on hard tasks it can actually solve, and to lack rigor in double-checking its own work. A rare exec comment on RL reward trade-offs ahead of release.
Related event: Musk Delays Grok 4.7, Cites Overly Harsh Length Penalties in RL(2 posts)→
More from Models
- User: Claude Opus spun for 3 hours on a bug, Grok 4.6 fixed it in 15 minutes — Daniel_Farinax · 2026-09-12
- After a week of full-time use, Google's Astra has no opinions on anything — lucasmeijer · 2026-09-12
- LMArena analyzed 30,086 answer pairs: different LLMs share just 43.1% of ideas — arena · 2026-09-12
- Four Reported Tricks Behind "Dumbed-Down" Models: Routing, Juice Cuts, Truncated Reasoning, MTP — vista8 · 2026-09-12
- GPT Astra users fret over 'honeymoon window': compute shortage rumors spark performance anxiety — TooManyB1tches · 2026-09-12
- tldraw founder lists 10 bugs in ChatGPT's new sketch feature — manosaie · 2026-09-12