GPT-6 Sol reportedly underperforms prior Sol on HealthBench, cyber, and bio/chem evals
sytelus · x · 2026-09-23
sytelus claims GPT-6 Sol underperforms the previous Sol on HealthBench, DeepSwe, cyber, and some bio/chem safety evaluations, suggesting possible safety regressions. The claim comes without a source and should be treated as unverified.
More from Models
- Abacus.AI CEO Bindu Reddy: SOL 6 Is a Real Regression, Losing to Terra Despite Better Pricing — bindureddy · 2026-09-23
- Opus 4.7 returns to Chat, Tasks and Claude Code; community extension keeps old models alive — repligate · 2026-09-23
- Arohan Desai's first take: Opus 5.5 feels like 4.5 at launch, less "Claudish" — _arohan_ · 2026-09-23
- Opus 5.5 released: beats Fable 5.1 in some tests, up to 40% cheaper than Opus 5 — soumitrashukla9 · 2026-09-23
- Alibaba's aggressive ASI roadmap: Qwen models to 5-10T params, 20GW by 2032 — bookwormengr · 2026-09-23
- Vals AI RSI Index: Opus 5.5 Beats Human Baseline on One of Five Tasks, Pulling RSI Estimate to July 2027 — ResultBackground2450 · 2026-09-23