Genentech's AutoSciBench auto-generates science agent benchmarks, cutting accuracy 25 points

Genentech · hf · 2026-10-07

Genentech introduced AutoSciBench, tackling benchmark saturation and costly manual construction for scientific agents. Tasks are represented as high-level concepts (domain, modality, reasoning type) plus low-level recipes (how question, environment, ground truth are built and verified). Agents attempt tasks; solver trajectories and judge feedback drive iterative recipe revision that closes shortcuts, and distilled refinement experience guides new concept generation.

Across computational biology, materials science, and clinical imaging, generated benchmarks reduce solver accuracy by 22.4 and 25.5 points in biology and materials respectively versus human-curated ones, with higher quality ratings—suggesting scientific-agent evaluation can adapt as capabilities advance.

Original post →

More from coding & agent

coding & agent channel →