Grok 4.6 ties GPT-5.6 Sol in engineering sciences benchmark

sanmikoyejo · x · 2026-08-28

Model performance varies by scientific domain. Top models from Anthropic and OpenAI dominate most fields, but in engineering sciences, SpaceX's Grok 4.6 ties GPT-5.6 Sol for second place (14.8%) at lower cost and token usage. Opus 5 leads in all domains except mathematical sciences, where Fable 5 (33.3%) and Sol (31.4%) take the lead.

Original post →

More from Models

Models channel →