Animation Bench tests 4 frontier models on rebuilding real web animations
himanshustwts · x · 2026-10-01
Physera launched Animation Bench, the first benchmark evaluating frontier coding agents on web animation reconstruction: 4 models, 48 tasks from 32 real commercial websites, 192 reconstructions, scored on visual similarity, motion consistency, and layout correctness.
Key finding: motion consistency is often incomplete or missing — current evals can't distinguish screenshot parity from shippable frontend reconstruction.
Leaderboard: GPT-6 Astra tops at 0.594 (motion only 0.473), Claude Fable 5.1 scores 0.548, GPT-6 Sol 0.516 ($0.45/task), Claude Opus 5.5 last at 0.507.
Related event: Animation Bench Launches to Test Web Animation Replication(4 posts)→
More from coding & agent
- Microsoft launches Power BI Authoring MCP server for natural-language semantic modeling — adnan_hashmi · 2026-10-01
- New notebook example shows how to use decision model Jev as an agent's next-step judge — _nerdai_ · 2026-10-01
- Swap in decision model Jev to end agent loops without extra LLM calls — _nerdai_ · 2026-10-01
- Agent harnesses show vast cost spreads: mini-swe hits best accuracy near direct-inference cost — xiye_nlp · 2026-10-01
- Hardware AI platform Flow raises $50M Series B at $750M valuation — ZehanWang · 2026-10-01
- autoharness turns your Claude Code sessions into a self-maintaining skills layer — dr_cintas · 2026-10-01