Loading...
正在加载...
请稍候

[论文] SCAFFOLD: A Large-Scale Structured Dataset of Computer Science Researc...

小凯 (C3P0) 2026年09月03日 00:46

论文概要

研究领域: CV
作者: Ranjit Raut, Aarav Subedi, Sagun Rai
发布时间: 2026-09-03
arXiv: 2509.00007

中文摘要

计算机科学论文严重依赖图表:架构图、系统流程图和管道示意图,它们通常比周围的文本携带更多信息。目前没有一个公开数据集将这类特定图表与标题、上下文、问题、答案和逐步推理配对,而这正是训练视觉语言模型理解它们所需要的。我们提出了SCAFFOLD,一个大规模结构化计算机科学研究图表数据集,包含图表问答和思维链推理轨迹。该数据集由来自arXiv计算机科学论文的(图像、标题、上下文、问答、思维链)元组组成,使用布局检测和PDF解析准备,并经过AI辅助的问题生成步骤。生成的大规模SCAFFOLD-157K数据集涵盖3,058篇论文的29,887张图表(157,387对),中等规模的SCAFFOLD-37K数据集(36,797对),以及小规模的SCAFFOLD-12K数据集(12,000对)。我们使用SCAFFOLD-12K对Qwen2.5-VL-3B-Instruct进行了基线实验。

原文摘要

Computer science papers rely heavily on diagrams: architecture drawings, system flowcharts, and pipeline schematics that often carry more information than the text around them. There is currently no public dataset that pairs this specific kind of figure with captions, context, questions, answers, and step-by-step reasoning, which is exactly what is needed to train a vision-language model to understand them. We present SCAFFOLD, a large-scale structured dataset of computer science research figures with diagram QA and Chain-of-Thought reasoning traces. This dataset consists of (image, caption, context, question-answer, chain-of-thought) tuples from arXiv computer science papers prepared using layout detection and PDF parsing, with an AI-assisted question-generation step. The resulting large-...


自动采集于 2026-09-03

#论文 #arXiv #CV #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录