Animation Bench full leaderboard: GPT-6 Astra leads as motion scores lag
himanshustwts · x · 2026-10-01
Physera published full details of Animation Bench: 4 frontier models, 48 tasks from 32 live commercial websites, 192 reconstructions.
- GPT-6 Astra — 0.594 overall (visual 0.710 / motion 0.473 / layout 0.628), 27 task wins, $3.04/task
- Claude Fable 5.1 — 0.548
- GPT-6 Sol — 0.516, cheapest at $0.45/task
- Claude Opus 5.5 — 0.507
Takeaway: models can mimic appearances but struggle to reconstruct page behavior — motion consistency is the weakest axis across the board. The team focuses on capability gaps in multimodal and commercial settings.
Related event: Animation Bench Launches to Test Web Animation Replication(4 posts)→
More from coding & agent
- LlamaIndex Draws ~600 in SF for Agent Document Processing Events — llama_index · 2026-10-01
- Prime Intellect on why enterprises should own their intelligence, not rent it — willcb · 2026-10-01
- Falcon Neo enters private beta with next-gen design infra and agent-friendly markup language — KadriJibraan · 2026-10-01
- AWS shows multi-account MCP pattern: shared AgentCore Gateway keeps data in each team's account — gethackteam · 2026-10-01
- OpenClaw Collapses Inter-Agent Receipts Into Single Expandable Rows — steipete · 2026-10-01
- Google replaces Gems with Skills, adopting Anthropic's agent-ready prompt standard — The Decoder · 2026-10-01