BIABench

A benchmark for evaluating AI agents on real-world, end-to-end bioimage analysis.

Results

Outcome is the mean over per-task means. Time and tokens are medians. * cost from billing records, otherwise token usage at list price. † wall-clock on a single NVIDIA A10 GPU.

Tasks

Citation

@misc{biabench2026,
  title         = {BIABench: Evaluating AI agents on real-world bioimage analysis tasks},
  author        = {Pan, Zixuan and Panzeri, Davide and Johanns, Lukas and Moor, Marilin
                   and Zhou, Yu and Peterson, Hedi and Shi, Yiyu and Chen, Jianxu},
  year          = {2026},
  eprint        = {2609.34274},
  archivePrefix = {arXiv},
  primaryClass  = {cs.AI},
  url           = {https://arxiv.org/abs/2609.34274}
}