BIABench: No AI agent scores above 0.19 on 3D bioimage analysis tasks
notredame · hf · 2026-10-02
Researchers released BIABench, an open benchmark testing whether AI agents can perform real-world bioimage analysis end to end.
- Tasks: 16 tasks reconstructed from peer-reviewed biological studies, preserving raw imaging data, scientific questions, and ground truth across 11 analysis subtasks and modalities from H&E histology to single-molecule localization microscopy.
- Scoring: each submission gets an outcome score (field-standard metrics vs. ground truth) and a process score (a VLM judges method choice and QC against expert-written rubrics).
- Findings: routine 2D tasks are solved well, but no agent scored above 0.19 on tasks adding a 3rd dimension or time axis. Biological specialization, stronger models, and expert instructions all failed to close the gap.
- Reliability: scores varied more between repeated runs of one agent than between agents, and without ground truth, correct runs couldn't be distinguished from wrong ones.
Released openly with data and code for evaluating—and eventually training—agents for long-horizon bioimage analysis.
More from coding & agent
- Voice-controlled AI agent unlocks developer's apartment on command — Baconbrix · 2026-10-02
- Pydantic hiring senior Rust engineer to build Monty, a sub-millisecond sandbox for AI agents — samuelcolvin · 2026-10-02
- Fabriq Developer launches: add 500+ integrations to your AI app with one prompt — ycombinator · 2026-10-02
- Companies are migrating from React Native to Swift as GenAI changes the calculus — mobileraj · 2026-10-02
- AI rewrites France's AROME weather model from scientific papers in four days — capetorch · 2026-10-02
- What should make you walk away from a voice-agent pilot? False confirmations — Legitimate-Pride-685 · 2026-10-02