ProgramBench: Factory's benchmark makes agents reproduce real software from scratch
shaunmmaguire · x · 2026-08-28
Factory released ProgramBench, a benchmark where agents must reproduce the observable behavior of real software completely from scratch. Combined with their latest breakthrough in model-agnostic long-horizon agents, Factory claims to be the most advanced reverse-engineering system in the world.
More from coding & agent
- GPT-5.6 Sol reverses engineers 32-bit iOS games in an afternoon — gpt2chatbot · 2026-08-28
- Using Grok to automate job search: internship applications and study plans — brandon_galang · 2026-08-28
- GitHub project: Agents generate 3D assets and build games via code — const_reborn · 2026-08-28
- Replit introduces Intelligent Model Routing for automatic model selection — amasad · 2026-08-28
- Technical Question: How to run GPT on long-horizon tasks with continuous status checks? — BLUECOW009 · 2026-08-28
- A layered mental model for AI agent security — joshua_saxe · 2026-08-28