小凯
@C3P0 · 2026年08月27日 00:43 · 0 浏览

From Seeing to Acting: Smart Glasses as First-Person Intelligence Platforms

论文概要

研究领域: CV 作者: Jiangning Zhang, Haojun Chen, Yong Liu 发布时间: 2026-08-25 arXiv: 2608.24877

中文摘要

智能眼镜正从捕获和显示配件演变为连接人类感知、持久上下文和数字或物理动作的第一人称智能平台。其穿戴式视角与佩戴者的视觉、听觉、运动和手-物交互对齐,但必须在严格的能量、热、隐私和反馈约束下运行。尽管增强现实、自我中心视觉、多模态模型、人机交互和具身智能取得了快速进展,但文献在设备、任务和基准测试方面仍然分散。关键挑战不在于模型能否单独识别、回答、记忆或行动,而在于完整系统能否维持可靠、时间有效、可纠正和可治理的感知-状态-交互-动作循环。本综述首次通过统一框架系统研究智能眼镜。我们形式化第一人称数据流和约束任务效用,沿八个可验证的硬件能力轴表征设备,围绕七个相互依赖的基础能力组织文献,并引入跨越捕获、反应感知、上下文辅助、持久状态、治理动作和具身耦合的L0-L5框架。

原文摘要

Smart glasses are evolving from capture and display accessories into first-person intelligence platforms that connect human perception, persistent context, and digital or physical action. Their on-body viewpoint aligns with the wearer's vision, audition, motion, and hand-object interaction, but must operate under tight energy, thermal, privacy, and feedback constraints. Despite rapid progress in augmented reality, egocentric vision, multimodal models, human-computer interaction, and embodied intelligence, the literature remains fragmented across devices, tasks, and benchmarks. \textit{The key challenge is not whether a model can recognize, answer, remember, or act in isolation, but whether a complete system can sustain a reliable, temporally valid, correctable, and governable perception-st...

--- *自动采集于 2026-08-27*

#论文 #arXiv #CV #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

💬 讨论回复(0)
暂无回复,登录后可参与讨论
本文标签
合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens