[论文] VIALS: A Benchmark for Visual Interpretation of Artifacts in the Life ...
论文概要
研究领域: ML 作者: Elaine Lau, Thanuka Udumulla, Lee Izhaki-Tavor, Francisco Guzmán, Nicholas Magazine, Jonas Mueller 发布时间: 2026-08-21 arXiv: 2608.21357
中文摘要
在专业的生命科学工作流程中,科学家日常需要解读视觉产物(凝胶印迹、显微镜图像、质粒图谱、流式细胞术图、分子结构等)以指导研究决策。我们推出了VIALS,一个包含161项此类解读任务的视觉问答基准,涵盖生物技术行业实验工作流程中检验的各类产物类型(而非来自出版物和教科书的精修图)。虽然前沿视觉语言模型现已能流利描述自然图像,但我们发现它们无法准确解读这些科学图像,反映出领域知识和领域特定视觉推理能力的局限。相比之下,具有相关领域专长的科学家认为这些视觉解读任务十分简单。无法类似解读此类图像的AI在专业生命科学工作流程中的效用将十分有限,因为这些产物是科学家推理、交流和做决策的核心。
原文摘要
In professional life sciences workflows, scientists routinely interpret visual artifacts (gel blots, microscopy images, plasmid maps, flow cytometry plots, molecular structures, ...) to inform research decisions. We introduce VIALS, a visual question-answering benchmark with 161 such interpretation tasks, spanning the types of artifacts examined throughout experimental workflows in the biotech industry (rather than polished figures from publications and textbooks). While frontier vision-language models can now fluently describe natural images, we find that they are unable to accurately interpret these scientific images, reflecting limitations in domain knowledge and domain-specific visual reasoning capabilities. In contrast, scientists with relevant domain expertise find these visual inter...
--- *自动采集于 2026-08-25*
#论文 #arXiv #ML #小凯