[论文] SCAFFOLD: A Large-Scale Structured Dataset of Computer Science Researc...

研究领域: CV 作者: Ranjit Raut, Aarav Subedi, Sagun Rai 发布时间: 2026-09-03 arXiv: 2509.00007

论文概要

研究领域: CV 作者: Ranjit Raut, Aarav Subedi, Sagun Rai 发布时间: 2026-09-03 arXiv: 2509.00007

中文摘要

计算机科学论文严重依赖图表:架构图、系统流程图和管道示意图,它们通常比周围的文本携带更多信息。目前没有一个公开数据集将这类特定图表与标题、上下文、问题、答案和逐步推理配对,而这正是训练视觉语言模型理解它们所需要的。我们提出了SCAFFOLD,一个大规模结构化计算机科学研究图表数据集,包含图表问答和思维链推理轨迹。该数据集由来自arXiv计算机科学论文的(图像、标题、上下文、问答、思维链)元组组成,使用布局检测和PDF解析准备,并经过AI辅助的问题生成步骤。生成的大规模SCAFFOLD-157K数据集涵盖3,058篇论文的29,887张图表(157,387对),中等规模的SCAFFOLD-37K数据集(36,797对),以及小规模的SCAFFOLD-12K数据集(12,000对)。我们使用SCAFFOLD-12K对Qwen2.5-VL-3B-Instruct进行了基线实验。

原文摘要

Computer science papers rely heavily on diagrams: architecture drawings, system flowcharts, and pipeline schematics that often carry more information than the text around them. There is currently no public dataset that pairs this specific kind of figure with captions, context, questions, answers, and step-by-step reasoning, which is exactly what is needed to train a vision-language model to understand them. We present SCAFFOLD, a large-scale structured dataset of computer science research figures with diagram QA and Chain-of-Thought reasoning traces. This dataset consists of (image, caption, context, question-answer, chain-of-thought) tuples from arXiv computer science papers prepared using layout detection and PDF parsing, with an AI-assisted question-generation step. The resulting large-...


*自动采集于 2026-09-03*

#论文 #arXiv #CV #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens