From Seeing to Acting: Smart Glasses as First-Person Intelligence Platforms
论文概要
研究领域: CV 作者: Jiangning Zhang, Haojun Chen, Yong Liu 发布时间: 2026-08-25 arXiv: 2608.24877
中文摘要
智能眼镜正从捕获和显示配件演变为连接人类感知、持久上下文和数字或物理动作的第一人称智能平台。其穿戴式视角与佩戴者的视觉、听觉、运动和手-物交互对齐,但必须在严格的能量、热、隐私和反馈约束下运行。尽管增强现实、自我中心视觉、多模态模型、人机交互和具身智能取得了快速进展,但文献在设备、任务和基准测试方面仍然分散。关键挑战不在于模型能否单独识别、回答、记忆或行动,而在于完整系统能否维持可靠、时间有效、可纠正和可治理的感知-状态-交互-动作循环。本综述首次通过统一框架系统研究智能眼镜。我们形式化第一人称数据流和约束任务效用,沿八个可验证的硬件能力轴表征设备,围绕七个相互依赖的基础能力组织文献,并引入跨越捕获、反应感知、上下文辅助、持久状态、治理动作和具身耦合的L0-L5框架。
原文摘要
Smart glasses are evolving from capture and display accessories into first-person intelligence platforms that connect human perception, persistent context, and digital or physical action. Their on-body viewpoint aligns with the wearer's vision, audition, motion, and hand-object interaction, but must operate under tight energy, thermal, privacy, and feedback constraints. Despite rapid progress in augmented reality, egocentric vision, multimodal models, human-computer interaction, and embodied intelligence, the literature remains fragmented across devices, tasks, and benchmarks. \textit{The key challenge is not whether a model can recognize, answer, remember, or act in isolation, but whether a complete system can sustain a reliable, temporally valid, correctable, and governable perception-st...
--- *自动采集于 2026-08-27*
#论文 #arXiv #CV #小凯