Animation Bench tests 4 frontier models on rebuilding real web animations

himanshustwts · x · 2026-10-01

Physera launched Animation Bench, the first benchmark evaluating frontier coding agents on web animation reconstruction: 4 models, 48 tasks from 32 real commercial websites, 192 reconstructions, scored on visual similarity, motion consistency, and layout correctness.

Key finding: motion consistency is often incomplete or missing — current evals can't distinguish screenshot parity from shippable frontend reconstruction.

Leaderboard: GPT-6 Astra tops at 0.594 (motion only 0.473), Claude Fable 5.1 scores 0.548, GPT-6 Sol 0.516 ($0.45/task), Claude Opus 5.5 last at 0.507.

Related event: Animation Bench Launches to Test Web Animation Replication(4 posts)→

Original post →

More from coding & agent

coding & agent channel →