论文概要
研究领域: ML
作者: Tica Lin, Deepak Chandran, Gauri Jagatap, Chen Chen, Andrea Fanelli, David Gunawan, Josh Kimball
发布时间: 2026-09-17
arXiv: 2609.20768
中文摘要
生成智能体越来越多地用于选择和叙述视频集锦,但它们通常基于非结构化或帧级表示运行。因此,观众难以验证其输出,也无法根据个人偏好进行引导。我们提出语义动作图(semantic action graph)——一种轻量级领域模式,将体育比赛表示为表演者、动作、接收者、时刻和状态节点,由角色、时间和结果边连接。该模式展示了三个关键特性:1) 连接的事件序列,2) 共享的封闭词汇表,3) 可寻址到帧的时刻——使其适合同时服务两类消费者:一个组合叙述集锦的智能体管线,以及一个让观众查询和检查同一结构的可视化界面。我们在 SportSAGE 中实例化了它——一个将四模块集锦管线与图界面配对的设计探索——并报告了来自12名足球迷的反馈。参与者对生成的集锦和叙述质量感到满意,并使用图界面搜索、导航和解读比赛集锦。这些结果提供了早期证据:一个小的、人类可读的模式可以同时支撑智能体生成和人类理解。
原文摘要
Generative agents are increasingly used to select and narrate video highlights, but they typically operate over unstructured or frame-level representations. Their output is consequently difficult for a viewer to verify and steer toward individual preferences. We present the semantic action graph, a lightweight domain schema that represents a sports match as performer, action, recipient, moment, and state nodes connected by role, temporal, and outcome edges. The schema demonstrates three key properties: 1) connected event sequences, 2) a shared, closed vocabulary, and 3) frame-addressable moments, making it suitable to serve two consumers at once: an agentic pipeline that composes narrated highlights, and a visual interface through which viewers query and inspect the same structure. We inst...
自动采集于 2026-09-19
#论文 #arXiv #ML #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。