GPT-6 Models Hallucinate Substantially Less Than Predecessors

ArtificialAnlys · x · 2026-09-23

Follow-up from Artificial Analysis: on the AA-Omniscience benchmark, both GPT-6 Sol and Luna hallucinate substantially less than their predecessors at max effort (Sol 92%→60%, Luna 93%→77%).

Related event: GPT-6 Sol divides reviewers: gains in coding agents, regression on DeepSWE, at half the price(11 posts)→

Original post →

More from Models

Models channel →