Same-Named Claude Opus 5 Behaved Like a Different Model Hours Apart in Side-by-Side Test

FishingCharming5604 · reddit · 2026-09-24

A Reddit user asked Claude Opus 5 the same 40 questions twice, hours apart, and found the second run looked things up 60% more often, wrote 50% more, and scored 90 vs 59 on source support. Opus 5.5 launched in between, but two other models barely changed.

The author concludes the likely cause was silent backend changes to settings around the model—thinking effort, available tools—rather than a model swap. Lesson: log every setting (including untouched ones) when benchmarking AI tools, or you may be measuring the settings, not the AI.

Original post →

More from Models

Models channel →