Pre-Registered Study: A 267-Word Spec Frame Cuts LLM Code Defects Across 5 Frontier Models

Sandeep Dhuri · hf · 2026-09-30

A pre-registered, five-model paired evaluation tested whether prepending a 267-word specification frame to prompts improves LLM-generated backend code. Across 50 realistic finance/healthcare/insurance tasks scored by 9 deterministic AST checkers, all five frontier models improved (mean defect reduction 0.16–0.70 per task, all Holm-adjusted sign tests significant); the frame arm won 95 of 100 differing comparisons and never made any model worse. Bandit found 53 medium-or-high issues in the bare arm vs 11 with the frame. All outputs, prompts, and pre-registration are published with a DOI.

Original post →

More from coding & agent

coding & agent channel →