PPTBench: a new benchmark testing if coding agents can rebuild visuals into editable slides
jiqizhixin · x · 2026-10-09
- Einsia AI's Navers Lab released PPTBench, evaluating coding agents on "visual coding": given a flowchart, web mockup, 3D scene sketch, or full PPT slide, the agent must rebuild it as runnable, editable code/slides.
- The challenge goes beyond making code run: agents must grasp objects, structure, spatial relations, and semantics from the image and encode them accurately — even GPT-6 Astra misreads flowcharts.
- Evaluation uses an Agentic Judge that starts from the delivered PPT and checks layer by layer whether the task was truly completed, without a fixed object tree, since reconstructions can look near-identical yet fail structurally.
More from coding & agent
- Google unveils universal Gemini agent for work: one prompt box for Q&A, knowledge work and code — sundarpichai · 2026-10-09
- Flipping the AI playbook: deploy agents on the client's own machine, then unplug — curious_vii · 2026-10-09
- Voyager: An Open Agent Harness Built for Creative Work, Not Coding — socialwithaayan · 2026-10-09
- Every's Marketer Has a Slack Agent Open PRs Straight from Figma — every · 2026-10-09
- Devs Ditch 500 .md Files: Rebuild AI Instructions From Scratch Each Model Release — bendee983 · 2026-10-09
- Catalyst raises $30M for autonomous AI trading agent, hiring engineers — jzlegion · 2026-10-09