Musk delays Grok 4.7 by days, says RL over-penalized response length

inductionheads · x · 2026-09-12

Elon Musk says Grok 4.7 needs a few more days to cook. He shared a training detail: the team may have over-penalized response length in RL, causing the model to give up on hard tasks it can actually solve, and to lack rigor in double-checking its own work. A rare exec comment on RL reward trade-offs ahead of release.

Related event: Musk Delays Grok 4.7, Cites Overly Harsh Length Penalties in RL(2 posts)→

Original post →

More from Models

Models channel →