Leaked V4 GA Scores Show Relaxed Safety Guardrails, Need for RL Advances
teortaxesTex · x · 2026-08-13
Scores from what appears to be V4 GA have surfaced. The model performs slightly weaker than K3 overall, but shows significantly fewer restrictions in traditionally sandbagged domains.
However, the performance gap suggests that scaling up is not enough; the developer needs another round of breakthroughs in reinforcement learning (RL) environments to close the gap with frontier models like 0731.
More from Models
- Grok 4.6 Released: Better Performance at Half the Price of Rivals — NinaDSchick · 2026-08-13
- DeepSeek V4 Pro Offers 10x Cheaper Cost Per Task Than GLM5.2 — ojasvi_yadav · 2026-08-13
- DeepSeek API Requests Timeout Amid Suspected Technical Issues — cedric_chee · 2026-08-13
- Free ChatGPT Users Silently Throttled for Overusing the 'Think' Button — Sauers_ · 2026-08-13
- Elon Musk Pushes Grok: Major Leap in Real-World Task Performance — elonmusk · 2026-08-13
- Hands-On Comparison: Grok 4.6 Edges Out DeepSeek in Success Rate, but Lags on Price — vista8 · 2026-08-13