PPTBench: 97% of AI-generated slides are valid files, only 2.57% defect-free
JaynitMakwana · x · 2026-09-29
Einsia releases PPTBench, a benchmark for visual coding via scientific diagram slide reconstruction: 500 tasks, 36 configurations, and 18,000 reconstructions.
Results expose current coding agents' weaknesses:
- Over 97% of runs produced valid slide files;
- But 64.03% failed basic semantic checks;
- Only 2.57% passed all gates with no recorded defect;
- Human visual inspection correlated with success at r = 0.88.
Key takeaway: generating code that compiles into a slide is easy; reconstructing the actual diagram logic is where models fail. Agents that can't inspect what they render are flying blind. Agent sessions are viewable on AgentGit.
More from coding & agent
- Garry Tan Backs AI Browsers Like AsideAI: Agent Password Management Is the Key Layer — garrytan · 2026-09-29
- Meeting Prep Agent Separates App State in SQLite from Long-Term Memory — Datrika_Nandhini · 2026-09-29
- Chestnut unveils 18-DOF Aero Hand plus matching exoskeleton for zero-gap humanoid data capture — chris_j_paxton · 2026-09-29
- RubyLLM hits 1.1M downloads in a month, 13M total — kieranklaassen · 2026-09-29
- Incident Memory Agent Turns Postmortems into Verified Triage Context via Hindsight — nagakeerthan · 2026-09-29
- InstaCloud launches agent-native serverless cloud that lets Claude Code and Cursor provision infra via one command — testingcatalog · 2026-09-29