SWE-bench Science benchmarks coding agents on scientific software repair

OpenMOSS-Team · hf · 2026-08-21

OpenMOSS-Team introduced SWE-bench Science, a benchmark designed to evaluate the ability of coding agents to resolve engineering tasks in scientific software. The benchmark reveals failure mechanisms in scientific software repair and the mixed effects of providing scientific guidance to agents.

Original post →

More from coding & agent

coding & agent channel →