Grok 4.6 ties GPT-5.6 Sol in engineering sciences benchmark
sanmikoyejo · x · 2026-08-28
Model performance varies by scientific domain. Top models from Anthropic and OpenAI dominate most fields, but in engineering sciences, SpaceX's Grok 4.6 ties GPT-5.6 Sol for second place (14.8%) at lower cost and token usage. Opus 5 leads in all domains except mathematical sciences, where Fable 5 (33.3%) and Sol (31.4%) take the lead.
More from Models
- Model comparison: Qwen, GLM, and Grok pass while ChatGPT and Claude fail — QuixiAI · 2026-08-28
- Qwen4 speculation: 27B model may see a leap in intelligence — Iory1998 · 2026-08-28
- Sparse attention can replace global attention without downsides — stochasticchasm · 2026-08-28
- llama.cpp Merges Qwen3.8-Flash-Next Support, GGUF Downloads Available — jacek2023 · 2026-08-28
- GPT-5.6 Sol matches Claude Fable 5 at 1/3 the cost — sanmikoyejo · 2026-08-28
- GLM-5.3-Flash retains 93% accuracy when quantized to 4-bit — amaarora · 2026-08-28