Paper Overview
- Field: Machine Learning
- Authors: Elaine Lau, Thanuka Udumulla, Lee Izhaki-Tavor, Francisco Guzmán, Nicholas Magazine, Jonas Mueller
- Published: 2026-08-21
- arXiv: 2608.21357
- VIALS contains 161 visual question-answering tasks based on real scientific artifacts from biotech industry experimental workflows.
- Artifact types include gel blots, microscopy images, plasmid maps, flow cytometry plots, and molecular structures.
- The benchmark deliberately avoids polished figures from publications and textbooks, favoring the messy, real-world images scientists actually analyze.
- Frontier vision-language models fail to accurately interpret these scientific images, despite strong performance on natural images.
- Failures reflect limitations in both domain knowledge and domain-specific visual reasoning.
- Domain-expert scientists find these tasks straightforward, highlighting a large human–AI gap.
- Implication: AI that cannot interpret such artifacts will have limited utility in professional life sciences work, where these visuals are central to reasoning, communication, and decisions.
Abstract (original)
In professional life sciences workflows, scientists routinely interpret visual artifacts (gel blots, microscopy images, plasmid maps, flow cytometry plots, molecular structures, ...) to inform research decisions. We introduce VIALS, a visual question-answering benchmark with 161 such interpretation tasks, spanning the types of artifacts examined throughout experimental workflows in the biotech industry (rather than polished figures from publications and textbooks). While frontier vision-language models can now fluently describe natural images, we find that they are unable to accurately interpret these scientific images, reflecting limitations in domain knowledge and domain-specific visual reasoning capabilities. In contrast, scientists with relevant domain expertise find these visual interpretation tasks easy.
Key Points
*Originally collected on 2026-08-25.*