Animation Bench tests coding agents on rebuilding web animations, and motion is their weak spot
himanshustwts · x · 2026-10-01
Physera released Animation Bench v1.0, a benchmark arguing that screenshot-parity evaluations can't distinguish visually plausible output from shippable frontend reconstruction. It asks multimodal coding agents to rebuild 48 real web animations from 32 live commercial sites (192 reconstructions), scoring visual, motion, and layout fidelity.
Leaderboard (0-1 reproduction score):
- GPT-6 Astra: 0.594, 27 task wins, $3.04/task
- Claude Fable 5.1: 0.548, 7 wins, $3.89/task
- GPT-6 Sol: 0.516, 8 wins, only $0.45/task
- Claude Opus 5.5: 0.507, 6 wins, $1.12/task
The standout finding: motion scores sit below 0.5 for every model (best: 0.473), while visual scores reach 0.63-0.71 — frontier agents can copy a page's look but still struggle to reproduce scroll, drag, and WebGL-driven behavior. Full task set and side-by-side reconstructions are in the appendix.
Related event: Animation Bench Launches to Test Web Animation Replication(4 posts)→
More from coding & agent
- Google AI proposes RRSI to stop recursive self-improving agents from overfitting benchmarks — burkov · 2026-10-01
- DevDay demo: Codex builds Minecraft for 30-year-old Game Boy hardware via ModRetro plugin — pvncher · 2026-10-01
- Full Talking Video From One Image in 30 Minutes for ~$5 — gorkem · 2026-10-01
- tldraw launches ChatGPT plugin: sketch your app's logic and Codex builds it — DavidKPiano · 2026-10-01
- Storing Claude's Memory Inside a Notes Repo With Auto-Commits — theshawwn · 2026-10-01
- Investment banker builds M-Terminal financial platform with Claude in 2 months of evenings — nikhil_pratap · 2026-10-01