Blender head-to-head: same prompt, 12 seconds, and one frontier model is in a different class
ZeroStateReflex · x · 2026-09-09
- Using the same Blender prompt for 12 seconds each, Claude Fable 5.1's output was clearly subpar compared to GPT-6 Astra — a gap that surprised a developer who builds AI tools inside Blender.
- The author's theory: the model isn't dumb, Blender simply isn't in Anthropic's training distribution yet.
- Analogy to Karpathy's chess observation: early models were terrible at chess until chess data entered the next training run, then improved immensely — the same will likely happen here.
- Takeaway: real in-software tests reveal capability gaps that generic benchmark scores hide. (Models not officially confirmed.)
Related event: Flagship Models Diverge Sharply in Same-Prompt Blender Test(2 posts)→
More from Fun
- Indie maker Marc Lou auctions ad spots on 10 of his muscles, bids start at $1,000 — Polymarket · 2026-09-09
- GPT-5.6 Sol writing unit tests for a five-line change, meme edition — JeremyNguyenPhD · 2026-09-09
- ChatGPT as an interior designer: demo in action — zeeg · 2026-09-09
- Another 'Lee Sedol moment': model pulls an AlphaGo-style move — polynoamial · 2026-09-09
- 'I have bunnies to feed': dev memes about AI subscription costs — tlakomy · 2026-09-09
- Gary Marcus mocks GPT-6 Astra video: OpenAI's 'human-level intelligence' claim is a lie — GaryMarcus · 2026-09-09