Arena Breaks Down False Attribution: GPT-6 Luna Rarely Misquotes but Misattributes 53% of the Time
arena · x · 2026-10-09
Arena's deep dive into False Attribution reveals varied failure patterns: GPT-6 Luna and Astra rarely misquote users (15.6% and 28.6%) but often misattribute statements to them (53.1% and 48.2%), while sibling model GPT-6 Sol has the highest rate of misstaging user history at 23.5%. Even models from the same lab fail in inconsistent ways.
More from Research
- Scientific ML is a loop: evaluation is an experiment on your whole modeling hypothesis — bravo_abad · 2026-10-09
- Models say no in chat but do it anyway: Simular reveals the agent safety gap — xwang_lk · 2026-10-09
- Three weeks, 19 lectures: a deep recap of Stanford AA203 from Euler equation to PPO — le_james94 · 2026-10-09
- Planning against a learned model seeks out exactly where the model errs flatteringly — le_james94 · 2026-10-09
- New cube packing record for n=12 at 2.9315 set with AI search method — CatAstro_Piyush · 2026-10-09
- Study: LLM judges of AI-scientist idea novelty are unreliable — MarioKrenn6240 · 2026-10-09