GamowLabs' LabBench: AI Agents Struggle to Pick the Next Best Wet-Lab Experiment

danielmckinn0n · x · 2026-09-30

GamowLabs, which recently opened a robotic biology lab for probing unstudied genetic variants, released LabBench, a public benchmark testing whether AI agents can plan and interpret wet-lab experiments. Built entirely from real, non-public lab data, it shows agents fail to commit to the next most-important experiment despite strong domain knowledge elsewhere in the eval. Blog and pre-print available.

Original post →

More from coding & agent

coding & agent channel →