Arena Breaks Down False Attribution: GPT-6 Luna Rarely Misquotes but Misattributes 53% of the Time

arena · x · 2026-10-09

Arena's deep dive into False Attribution reveals varied failure patterns: GPT-6 Luna and Astra rarely misquote users (15.6% and 28.6%) but often misattribute statements to them (53.1% and 48.2%), while sibling model GPT-6 Sol has the highest rate of misstaging user history at 23.5%. Even models from the same lab fail in inconsistent ways.

Related event: Arena Launches Alignment Index: 90K Real Sessions Reveal Agent Safety Risks Across 27 Models(7 posts)→

Original post →

More from Research

Research channel →