论文概要
研究领域: ML
作者: Muzhe Wu, Zuchen Li, Xu Wang, Anhong Guo
发布时间: 2026-09-21
arXiv: 2609.24955
中文摘要
物理任务的视觉指导通常在一种情境下制作,却在另一种情境下使用,这要求用户将演示中的工具、材料和空间关系转化到自己的环境中。我们提出「生成式教程」(Generative Tutorial),一个用于实时视觉指导的概念框架,在用户的环境和任务流程中展示预期结果和操作方法。通过对 15 个物理任务的最先进图像和视频生成模型进行形成性评估,我们识别了失败模式和潜在收益。基于这些发现,我们构建了一个增强现实原型系统,利用观察到的操作空间上下文和先前动作的预测视觉结果,主动生成目标图像和演示视频。一项 24 人参与的实验室研究发现,与预制作的指导相比,使用该系统时任务完成质量更高、感知操作空间对应度更强、步骤确认间隔更短。定性发现强调了情境相似性如何影响信任、生成错误如何影响理解,以及指导交付应如何适应用户需求,为未来设计提供了启示。
原文摘要
Visual instructions for physical tasks are typically authored in one context and followed in another, requiring users to translate demonstrated tools, materials, and spatial relationships into their own environment. We introduce Generative Tutorial, a conceptual framework for live visual instruction that depicts intended outcomes and actions within the user's environment and task flow. A formative evaluation of state-of-the-art image and video generation identifies failures and potential benefits across 15 physical tasks. Drawing on these findings, we build an augmented-reality prototype system that proactively generates goal images and demonstration videos using observed workspace context and predicted visual outcomes of preceding actions. A 24-participant lab study found higher task perf...
自动采集于 2026-09-23
#论文 #arXiv #ML #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。