Genentech's AutoSciBench auto-generates science agent benchmarks, cutting accuracy 25 points
Genentech · hf · 2026-10-07
Genentech introduced AutoSciBench, tackling benchmark saturation and costly manual construction for scientific agents. Tasks are represented as high-level concepts (domain, modality, reasoning type) plus low-level recipes (how question, environment, ground truth are built and verified). Agents attempt tasks; solver trajectories and judge feedback drive iterative recipe revision that closes shortcuts, and distilled refinement experience guides new concept generation.
Across computational biology, materials science, and clinical imaging, generated benchmarks reduce solver accuracy by 22.4 and 25.5 points in biology and materials respectively versus human-curated ones, with higher quality ratings—suggesting scientific-agent evaluation can adapt as capabilities advance.
More from coding & agent
- Cursor adds Cloud Agents API endpoints for environment builds with status and error codes — tetsuoai · 2026-10-07
- Figure CEO: filling Vietnam's brutal visa form was our AGI test — now an agent passed it — adcock_brett · 2026-10-07
- Vite+ 1.1 released: 24% faster vp dev startup, 30% less memory, clearer prompts — irvinebroque · 2026-10-07
- solid-yield Brings Generator-Based Type-Safe Components to Solid 2 as AI Sparks a Yield Renaissance — samgoodwin89 · 2026-10-07
- Rex, a Coding Agent Multiplexer, Opens Mailing List Invites for Early Testing — DanielLockyer · 2026-10-07
- Anaconda Combines Coding Agent Swarms and AI Attack Testing After Three Acquisitions — shashib · 2026-10-07