Grok 4.7 delayed: Musk says RL may have penalized response length too much

Scobleizer · x · 2026-09-12

xAI is holding back Grok 4.7 for "a few more days." Elon Musk explained the model still gives up too early on hard tasks it can actually do and isn't rigorous enough checking its own work — possibly because RL penalized response length too much. Observers call the public postmortem a notable admission of a training tradeoff gone wrong.

Related event: Musk Delays Grok 4.7, Cites Overly Harsh Length Penalties in RL(2 posts)→

Original post →

More from Models

Models channel →