Musk Delays Grok 4.7, Cites Overly Harsh Length Penalties in RL
Elon Musk said Grok 4.7 needs a few more days, revealing that overly harsh RL penalties on response length may have caused the model to give up prematurely on tasks it could actually solve.
2026-09-12 ~ 2026-09-12 · 2 related posts
- Musk delays Grok 4.7 by days, says RL over-penalized response length — inductionheads · 2026-09-12
1 near-duplicate retellings: Scobleizer