Musk Delays Grok 4.7, Cites Overly Harsh Length Penalties in RL

Elon Musk said Grok 4.7 needs a few more days, revealing that overly harsh RL penalties on response length may have caused the model to give up prematurely on tasks it could actually solve.

2026-09-12 ~ 2026-09-12 · 2 related posts

1 near-duplicate retellings: Scobleizer