← 首页|学术|From Seeing to Acting: Smart Glasses as First-Person Intelligence Platforms
cs.CV · 2608.24877 · 2026-08-25

From Seeing to Acting: Smart Glasses as First-Person Intelligence Platforms

Jiangning Zhang, Haojun Chen, Yong Liu
Smart GlassesEgocentric AISurvey
💬 首个用统一框架系统研究智能眼镜的综述:提出 L0-L5 能力分级(捕获→反应式感知→情境辅助→持久状态→受治理动作→具身耦合),并给出可验证的部署与评估协议。

🎯 背景

智能眼镜正从单纯的拍摄/显示配件演化为连接人类感知、持久上下文、数字或物理动作的第一人称智能平台。其贴身视角与佩戴者的视觉、听觉、运动和手物交互天然对齐,但必须在严格的能耗、散热、隐私和反馈约束下运行。尽管增强现实、第一人称视觉、多模态模型、人机交互和具身智能领域都在快速进展,相关文献目前仍在设备、任务和基准之间高度碎片化。

🔬 方法

这是首个用统一框架系统研究智能眼镜的综述。作者形式化了第一人称数据流和受限任务效用,沿 8 个可验证的硬件能力轴刻画设备,围绕 7 项相互依赖的基础能力组织现有文献,并引入一个跨越捕获、反应式感知、情境辅助、持久状态、受治理动作、具身耦合的 L0-L5 分级框架。论文进一步跨 9 个应用场景连接任务、数据集、系统、产品、利益相关者、失败后果与证据缺口,提出一个九维部署框架、一个按声明分级的评估协议,以及一个从受控测量到纵向现场验证与审计的证据阶梯。

📊 结果

作为综述论文,产出是一套分类框架、评估协议和研究路线图,而非实验数字结果,目标是让智能眼镜系统之间的比较、部署评估和可复现研究变得更容易。

原文摘要

▶ Abstract
Smart glasses are evolving from capture and display accessories into first-person intelligence platforms that connect human perception, persistent context, and digital or physical action. Their on-body viewpoint aligns with the wearer's vision, audition, motion, and hand-object interaction, but must operate under tight energy, thermal, privacy, and feedback constraints. Despite rapid progress in augmented reality, egocentric vision, multimodal models, human-computer interaction, and embodied intelligence, the literature remains fragmented across devices, tasks, and benchmarks. \textit{The key challenge is not whether a model can recognize, answer, remember, or act in isolation, but whether a complete system can sustain a reliable, temporally valid, correctable, and governable perception-state-interaction-action loop.} This survey is \textit{the \textbf{first} to systematically study smart glasses through such a unified framework}. We formalize first-person data flow and constrained task utility, characterize devices along eight verifiable hardware capability axes, organize the literature around seven interdependent foundational capabilities, and introduce an L0-L5 framework spanning capture, reactive perception, contextual assistance, persistent state, governed action, and embodied coupling. Across nine application scenes, we connect tasks with datasets, systems, products, stakeholders, failure consequences, and evidence gaps. We further present a nine-dimensional deployment framework, a claim-conditioned evaluation protocol, and an evidence ladder from controlled measurement to longitudinal field validation and audit. Together, these elements make smart glasses more comparable, deployable, and reproducibly evaluated, while outlining a roadmap toward trustworthy first-person intelligence.
来源: arXiv:2608.24877 · 精读由高松灯生成,基于摘要与 arXiv 页面信息