Results
Outcome is the mean over per-task means. Time and tokens are medians. * cost from billing records, otherwise token usage at list price. † wall-clock on a single NVIDIA A10 GPU.
The paper’s configurations under the brief instruction, three runs per agent–task pair. Cost per run on a log scale. Filled points are the frontier: nothing is both cheaper and better. Hover for values.
01refused: all three runs blocked by the provider's safety filter
The six agents on GPT-5.6 Sol under the brief instruction. Hover a value for its three runs.
Tasks
Citation
@misc{biabench2026,
title = {BIABench: Evaluating AI agents on real-world bioimage analysis tasks},
author = {Pan, Zixuan and Panzeri, Davide and Johanns, Lukas and Moor, Marilin
and Zhou, Yu and Peterson, Hedi and Shi, Yiyu and Chen, Jianxu},
year = {2026},
eprint = {2609.34274},
archivePrefix = {arXiv},
primaryClass = {cs.AI},
url = {https://arxiv.org/abs/2609.34274}
}