[论文] Semantic Action Graph: A Shared Representation for Agent Grounding and...

研究领域: ML 作者: Tica Lin, Deepak Chandran, Gauri Jagatap, Chen Chen, Andrea Fanelli, David Gunawan, Josh Kimball 发布时间: 2026-09-17 arXiv: 2609.20768

论文概要

研究领域: ML 作者: Tica Lin, Deepak Chandran, Gauri Jagatap, Chen Chen, Andrea Fanelli, David Gunawan, Josh Kimball 发布时间: 2026-09-17 arXiv: 2609.20768

中文摘要

生成智能体越来越多地用于选择和叙述视频集锦,但它们通常基于非结构化或帧级表示运行。因此,观众难以验证其输出,也无法根据个人偏好进行引导。我们提出语义动作图(semantic action graph)——一种轻量级领域模式,将体育比赛表示为表演者、动作、接收者、时刻和状态节点,由角色、时间和结果边连接。该模式展示了三个关键特性:1) 连接的事件序列,2) 共享的封闭词汇表,3) 可寻址到帧的时刻——使其适合同时服务两类消费者:一个组合叙述集锦的智能体管线,以及一个让观众查询和检查同一结构的可视化界面。我们在 SportSAGE 中实例化了它——一个将四模块集锦管线与图界面配对的设计探索——并报告了来自12名足球迷的反馈。参与者对生成的集锦和叙述质量感到满意,并使用图界面搜索、导航和解读比赛集锦。这些结果提供了早期证据:一个小的、人类可读的模式可以同时支撑智能体生成和人类理解。

原文摘要

Generative agents are increasingly used to select and narrate video highlights, but they typically operate over unstructured or frame-level representations. Their output is consequently difficult for a viewer to verify and steer toward individual preferences. We present the semantic action graph, a lightweight domain schema that represents a sports match as performer, action, recipient, moment, and state nodes connected by role, temporal, and outcome edges. The schema demonstrates three key properties: 1) connected event sequences, 2) a shared, closed vocabulary, and 3) frame-addressable moments, making it suitable to serve two consumers at once: an agentic pipeline that composes narrated highlights, and a visual interface through which viewers query and inspect the same structure. We inst...


*自动采集于 2026-09-19*

#论文 #arXiv #ML #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens