Benchling benchmarks Claude and ChatGPT on wet-lab protocols: helpful, not solved
nlarusstone · x · 2026-09-22
Benchling's Head of AI Nicholas Larus-Stone discusses the Bench-Bench study on the Ion Genomics podcast: leading LLMs like Claude and ChatGPT were tasked with redesigning wet-lab protocols and did fine but didn't impress. Key points:
- Much of good benchwork lives in researchers' heads, not papers or protocols—why LLMs still struggle with wet-lab reasoning
- He believes pharma companies should be spending far more on AI usage than they do
- Also covers why benchmarking LLMs matters, open-source vs proprietary models, and founding Bits in Bio
Other referenced benchmarks include BiomniBench and PromptBio-Bench.
More from Apps
- Flight booking is the worst example for AI agents, argues developer willcb — willcb · 2026-09-22
- Meta Muse first test: sorts a year of expenses across Gmails in 15 minutes — armand_ruiz · 2026-09-22
- Meta Muse rally: stock up 11% as Morgan Stanley pegs $1B revenue per 100M users — firstadopter · 2026-09-22
- Hands-on with Pexo: conversational video agent makes explainer films, but credits burn fast — sujingshen · 2026-09-22
- Why Instagram data may be Muse's ace over ChatGPT, Claude and even Gemini — brandon_galang · 2026-09-22
- Validate a startup idea before building it: indie hacker makes an AI promo video with Pexo — songguoxiansen · 2026-09-22