Grok 4.5 Performs Well in Security Evaluation
zeeg · x · 2026-07-11
The author stated that while he hasn't yet run the same tests using GPT-5.6, he believes the difference between 5.6 and 5.5 might not be that critical, and making direct comparisons is harder now. He is currently still using Sonnet 4.6 but is considering switching to Grok due to the price/accuracy trade-off.\n\nThe original post mentioned: Grok 4.5 performed very well on the Warden security review benchmark. Although it has some speed issues, its cost and accuracy are solid; the author noted that its performance surpassed all the Opus models he has tested.
Related event: Grok 4.5 Evaluations: Strong Cost-Performance in Mid-to-High Budgets(5 posts)→
More from Models
- DeepMind-Princeton paper shows LLMs causally use confidence to decide whether to answer — GoogleDeepMind · 2026-09-07
- Philosopher asks GPT-6 to review his Oxford book: result rivals top-journal reviews — anselm · 2026-09-07
- Leaker claims xAI is preparing Grok 4.7, hints at another surprise — mark_k · 2026-09-07
- Local LLMs now near Opus-level — what's still keeping them behind closed models? — mrsalvadordali · 2026-09-07
- Blind test of 12 models finds Fable 5.1 reads least like AI at 14%, Gemini 3.8 Flash worst at 77% — PawelHuryn · 2026-09-07
- Users miss the old Claude that used emojis: newer versions turn oddly poetic — JoshuaJosephson · 2026-09-07