Grok 4.7 delayed: Musk says RL may have penalized response length too much
Scobleizer · x · 2026-09-12
xAI is holding back Grok 4.7 for "a few more days." Elon Musk explained the model still gives up too early on hard tasks it can actually do and isn't rigorous enough checking its own work — possibly because RL penalized response length too much. Observers call the public postmortem a notable admission of a training tradeoff gone wrong.
Related event: Musk Delays Grok 4.7, Cites Overly Harsh Length Penalties in RL(2 posts)→
More from Models
- GPT-6 Astra's chain-of-thought can evade monitoring; auditability may be the next differentiator — buckymoore · 2026-09-12
- OpenAI image generator's output format is unpredictable, users prefer the previews — No_Body_4834 · 2026-09-12
- Users report Gemini Pro 3.1 refusing even simple requests — kostaslamprou · 2026-09-12
- Startup reportedly builds autonomous drone system using GPT-6 Astra to track people from a single image — Polymarket · 2026-09-12
- Nex-N2.5 Pro, a 397B multimodal model focused on Computer Use, quietly lands on OpenRouter — nikola_mr64990 · 2026-09-12
- GPT-6 fails to improve on molecular property prediction, fueling AGI skepticism — GaryMarcus · 2026-09-12