BIABench: No AI agent scores above 0.19 on 3D bioimage analysis tasks

notredame · hf · 2026-10-02

Researchers released BIABench, an open benchmark testing whether AI agents can perform real-world bioimage analysis end to end.

Released openly with data and code for evaluating—and eventually training—agents for long-horizon bioimage analysis.

Original post →

More from coding & agent

coding & agent channel →