Pre-Registered Study: A 267-Word Spec Frame Cuts LLM Code Defects Across 5 Frontier Models
Sandeep Dhuri · hf · 2026-09-30
A pre-registered, five-model paired evaluation tested whether prepending a 267-word specification frame to prompts improves LLM-generated backend code. Across 50 realistic finance/healthcare/insurance tasks scored by 9 deterministic AST checkers, all five frontier models improved (mean defect reduction 0.16–0.70 per task, all Holm-adjusted sign tests significant); the frame arm won 95 of 100 differing comparisons and never made any model worse. Bandit found 53 medium-or-high issues in the bare arm vs 11 with the frame. All outputs, prompts, and pre-registration are published with a DOI.
More from coding & agent
- Denying GPT-6.1 Sol write access as orchestrator cut coding costs 77%, at 6x runtime — GapNew4766 · 2026-09-30
- 8 research agents self-train a 30B model for 144 hours in RSIArena livestream experiment — my_cat_can_code · 2026-09-30
- OpenAI's Nan Yu: a boss agent running other agents is just one agent with extra steps — victor_explore · 2026-09-30
- Open-source MCP server lets agents query 124.6B TikTok data points — operatorarkay · 2026-09-30
- Dev argues truly always-on autonomous agents have never actually been tried — jacob_posel · 2026-09-30
- Skill scaffold template: same shape, faster review, fewer broken CIs — blaizedsouza · 2026-09-30