GPT-6 Models Hallucinate Substantially Less Than Predecessors
ArtificialAnlys · x · 2026-09-23
Follow-up from Artificial Analysis: on the AA-Omniscience benchmark, both GPT-6 Sol and Luna hallucinate substantially less than their predecessors at max effort (Sol 92%→60%, Luna 93%→77%).
More from Models
- Five Hours of Heavy Opus 5.5 Use on /medium Burned Only 4% of the Weekly Limit — rudrank · 2026-09-23
- GPT-6-Luna vs GPT-5.6-Luna: A Voxel Pagoda Generation Showdown — BLUECOW009 · 2026-09-23
- China Probes DeepSeek and Moonshot Over Data Breaches, The Information Reports — teortaxesTex · 2026-09-23
- Opus 5.5 tops AI index at 58, undercuts GPT-6 Astra by 60% but burns 4x more tokens — johnseach · 2026-09-23
- Altman clarifies: OpenAI's new release is voice, not video, 'sorry to disappoint' — sama · 2026-09-23
- Grok 4.7, Opus 5.5, and GPT-6 Sol/Luna all shipped in one insane September week — altryne · 2026-09-23