Flagship Models Diverge Sharply in Same-Prompt Blender Test

A developer's same-prompt Blender test found GPT-6 Astra's 12-second output clearly superior to Claude Fable 5.1, attributing the gap mainly to training data distribution: models can't generate what their training data never covered.

2026-09-08 ~ 2026-09-09 · 2 related posts