Rules to Tools: Executable Checks Beat Text for SciCode Agent Repairs 29/30 vs 26/30
AbhiVakil29112001 · hf · 2026-10-02
Rules to Tools (R2T) converts public scientific requirements into prepared executable checks for LLM coding agents, compared against equivalent written requirements in matched SciCode repair groups. Tool groups reach 29/30 complete repairs vs 26/30 with text across two task-ID cohorts, though a larger shared-definition cohort ties at 13/24. In a matched PDE comparison, checks score 24/24 vs text's 23/24 with 31.2% lower reported model output. Gains are task-dependent, and public CPU use rises in both cohorts.
More from coding & agent
- Dev recreates a DOOM-like game with GPT, nailing the original's gory feel — DeryaTR_ · 2026-10-02
- Giving an AI agent real control of a homelab — responsibly — Academic_Wolverine · 2026-10-02
- Process-mining agents found 20 steps and 7 loops in a workflow documented as 7 steps — vasuman · 2026-10-02
- Geoffrey Huntley: Forget reading code—your verification properties are all that matters — kieranklaassen · 2026-10-02
- Claude Skills explained: why prewritten PDF scripts beat pasting prompts every time — lxfater · 2026-10-02
- Eye.Art Polyphemus: a chat-first MCP for image generation and reference-based edits — axiomofaxiom · 2026-10-02