PPTBench: 97% of AI-generated slides are valid files, only 2.57% defect-free

JaynitMakwana · x · 2026-09-29

Einsia releases PPTBench, a benchmark for visual coding via scientific diagram slide reconstruction: 500 tasks, 36 configurations, and 18,000 reconstructions.

Results expose current coding agents' weaknesses:

Key takeaway: generating code that compiles into a slide is easy; reconstructing the actual diagram logic is where models fail. Agents that can't inspect what they render are flying blind. Agent sessions are viewable on AgentGit.

Original post →

More from coding & agent

coding & agent channel →