BIABench: AI agents ace routine 2D bioimage analysis but cap at 0.19 on 3D tasks

BIABench: Evaluating AI agents on real-world bioimage analysis tasks

Zixuan Pan, Davide Panzeri, Lukas Johanns, Marilin Moor, Yu Zhou, Hedi Peterson, Yiyu Shi, Jianxu Chen

cs.AI, cs.CL

2026-09-28

BIABench turns 16 published studies into end-to-end bioimage tasks: routine 2D analyses reach 0.85–0.96, the hardest 3D tasks cap at 0.25, and run-to-run variance is 6x the agent choice.

What problem this solves

The daily image analysis of a biology lab, counting stained cells, segmenting nuclei, tracking divisions, measuring colocalization, needs both specialist software (Fiji, Cellpose) and someone who can code. AI agents look like a fix, but no benchmark had tested whether one can carry a real analysis from raw images to a scientific conclusion. Bioimage benchmarks grade single-step models one step at a time; agent benchmarks such as BixBench3 and BiomniBench test text QA or tabular omics, and none of them touch microscopy.

Two things make real analyses hard. The data are huge: 13.9 GB of raw input across 16 tasks, including 3D volumes and 800-frame movies, far too large for context, so an agent must work through code, specialized software and rendered views. The task is also open-ended: no prescribed pipeline, the agent picks the tools (GUI software included) and must deliver masks, tracks and statistics.

Method

BIABench (Notre Dame, ISAS Dortmund, University of Tartu) rebuilds tasks from published biological studies: keep the raw images and the scientific question, treat the peer-reviewed result as ground truth. The recipe can keep turning new papers into new tasks, and instructions forbid retrieving the source publication.

Scoring has two layers. An outcome score compares the required output files against ground truth with field-standard metrics (Dice, the Cell Tracking Challenge TRA measure, the KS statistic); missing files score zero, and rankings use this score alone. A process score has a VLM judge method choice and quality control against expert rubrics; it agrees with a human expert on 87% of items (Cohen's κ = 0.67). All six agents run through one shared interface with files recovered from the filesystem, so the harness does not confound the model comparison:

Every configuration-task pair runs three times under a wall-clock budget (4 h, 2 h in most sessions). Instructions come at two levels: a brief biologist's description, and a detailed expert protocol with a recommended pipeline and parameter hints.

Results

First with the model held fixed (GPT-5.6 Sol), then swapping models behind one harness:

ConfigurationMetricResult
Claude Code (GPT-5.6 Sol)Mean outcome over 16 tasks0.65, best of six; CopilotJ lowest at 0.50
Routine 2D tasks (NF-κB, HeLa segmentation, counting)Best agent0.85–0.96; every agent at least 0.56
5D nuclear-pore kinetics / 3D puncta quantificationCap over all model × instruction combos≤0.25; best in the GPT-5.6 study 0.19 / 0.05
Variance decompositionRun-to-run vs choice of agent21% vs 3%, a 6× gap; the task itself 63%
DeepSeek-V4-FlashCost efficiency80% of the best score for 2% of the cost ($0.09 vs $4.53 per run)

Three findings hold. Adding a dimension or a time axis collapses performance: on nuclear-pore kinetics and 3D puncta quantification, neither stronger models nor detailed instructions got past 0.25. Biological specialization buys nothing: general-purpose agents took or tied the top mean on 14 of 16 tasks and led by 0.32 on bacterial tracking. Agents are unstable: 40% of agent-task pairs swung more than 0.2 across three identical runs, and taking the best of three squeezed five of six agents within 0.04 of one another, so the gap is reliability, not peak capability.

The detailed expert protocol moved the mean by +0.01 and mostly redistributed scores: bacterial tracking rose from 0.32 to 0.94 while cell counting fell from 0.83 to 0.42.

Self-checking is the deeper failure. Process and outcome scores correlate at r = 0.29 (0.09 after excluding near-zero outcomes), and expert-assigned process scores correlated at -0.03; Agentic-J had the highest process score (0.87) with a mean outcome of only 0.55. Runs scoring below 0.1 took twice as long (median 21 vs 10 minutes), and many closed by claiming completion (9 of 9 for DeepSeek Harness on V4-Pro, 3 of 3 for Codex). Provider safety filters blocked Claude Code and Codex entirely on the SARS-CoV-2 colocalization task.

Why it matters

For agent builders the message is blunt: the bottleneck is long-horizon planning and self-checking, not domain knowledge, and expert protocols do not rescue it. Run-to-run variance is six times the variance from the choice of agent, so single-run leaderboards prove little in this domain; evaluations must repeat runs. For the bioimage community this is the first end-to-end benchmark, with tasks, rubrics and code released openly (Hugging Face and GitHub), usable for evaluation and eventually for training. The cost result is practical too: at $0.09 per run, DeepSeek-V4-Flash reaches 80% of the best score as long as the work stays 2D.

Limitations

The authors' own list: tasks are scored at the study's endpoint, which assumes human experts would score near-perfectly, and current agents sit far from that; several biology-specific agents were built for human-in-the-loop use and underperform when run autonomously, which softens the "specialization is useless" reading; and the team co-developed Agentic-J, one of the evaluated agents (disclosed, and it did not place highly).

Questions that remain: the VLM judge is lenient, passing 27% of the items the expert failed and almost never scoring below 0.5, so its discriminative power is limited; 16 tasks with three runs per configuration is thin for statistics; GLM-5.1 and V4-Flash could not run under Claude Code because empty API responses were logged as finished turns, leaving the model × harness matrix incomplete; and most sessions used a 2-hour wall-clock budget with truncated runs scored on whatever files existed, which penalizes slower agents such as Agentic-J (median 31.7 minutes).

Terms

Source

Related papers

All paper explainers