Blogger: OpenAI models crush Claude Opus at scientific automation, not just benchmarks
VraserX · x · 2026-09-30
Blogger VraserX argues OpenAI models dominate math, science, computer use and 3D/world modeling, showing they're built for serious scientific work rather than coding benchmarks. He concedes Claude Opus is better at coding, but claims even GPT-6.1 Sol, a smaller distilled model, crushes it at automating real science — a gap Tech Twitter underappreciates.
More from Models
- 32 researchers spent 8 months on the most comprehensive survey of tokenization ever — christopher · 2026-10-01
- GPT-6.1 Sol reportedly tops MathArena: 86.3% accuracy at $0.94 vs Astra's 81.9% at $2.26 — 141_1337 · 2026-10-01
- Defending Astra's Caution: An Agent That Stops to Ask Is Doing It Right — brandon_galang · 2026-10-01
- Dev complains Opus 5.5 burned 80% of quota in 3 days, switching models — MaziyarPanahi · 2026-10-01
- Liquid AI's LongevityBench: compact LFMs beat frontier models on several aging tasks — JosephJacks_ · 2026-10-01
- GPT 6.1 Astra ultra code mode builds a three.js ocean scene — OpenAIDevs · 2026-09-30