Rules to Tools: Executable Checks Beat Text for SciCode Agent Repairs 29/30 vs 26/30

AbhiVakil29112001 · hf · 2026-10-02

Rules to Tools (R2T) converts public scientific requirements into prepared executable checks for LLM coding agents, compared against equivalent written requirements in matched SciCode repair groups. Tool groups reach 29/30 complete repairs vs 26/30 with text across two task-ID cohorts, though a larger shared-definition cohort ties at 13/24. In a matched PDE comparison, checks score 24/24 vs text's 23/24 with 31.2% lower reported model output. Gains are task-dependent, and public CPU use rises in both cohorts.

Original post →

More from coding & agent

coding & agent channel →