GPT-6 Sol reportedly underperforms prior Sol on HealthBench, cyber, and bio/chem evals

sytelus · x · 2026-09-23

sytelus claims GPT-6 Sol underperforms the previous Sol on HealthBench, DeepSwe, cyber, and some bio/chem safety evaluations, suggesting possible safety regressions. The claim comes without a source and should be treated as unverified.

Related event: GPT-6 Sol/Luna reviews: half price, fewer hallucinations, but regressions on some benchmarks(15 posts)→

Original post →

More from Models

Models channel →