ProgramBench finds no model can fully rebuild software from scratch

zetalyrae · x · 2026-07-25

ProgramBench tests whether models can rebuild software from scratch

A screenshot shows the paper “ProgramBench: Can Language Models Rebuild Programs From Scratch?” and notes that hashcards was included in the study.

The paper’s premise is that current benchmarks often test narrow tasks like fixing a single bug or adding one feature, but real software work requires something broader: designing and implementing an entire codebase from a program plus documentation.

Key points from the paper shown in the image:

The image also lists eudoxia0/hashcards as a plain-text spaced repetition flashcard system, suggesting it was one of the software targets used in the benchmark.

Original post →

More from coding & agent

coding & agent channel →