GamowLabs' LabBench: AI Agents Struggle to Pick the Next Best Wet-Lab Experiment
danielmckinn0n · x · 2026-09-30
GamowLabs, which recently opened a robotic biology lab for probing unstudied genetic variants, released LabBench, a public benchmark testing whether AI agents can plan and interpret wet-lab experiments. Built entirely from real, non-public lab data, it shows agents fail to commit to the next most-important experiment despite strong domain knowledge elsewhere in the eval. Blog and pre-print available.
More from coding & agent
- Dev switches to local Codex to keep cowork data on-premise — HanchungLee · 2026-09-30
- Whiteboarding how Microsoft Work IQ + Copilot assembles enterprise context for LLMs — clamanna · 2026-09-30
- What DevDay 2026 Reveals About OpenAI's Cloud Agent Ambitions and Who Gets Access — TheTuringPost · 2026-09-30
- "Vibe coding is the new Visual Basic" — the AI coding debate in one meme — generativist · 2026-09-30
- Browser Use Launches Ultrafast: 10x Faster Web Agents That Compare Flights for $0.004 — garrytan · 2026-09-30
- OpenAI DevDay: 20+ Updates Including Dots Agents, GPT-6.1 Sol, Ultrafast and $500 Pro 500 Plan — btibor91 · 2026-09-30