← 首页|学术|UrbanGround: Spatial Agency in Real-Scale City
cs.CV · 2608.27456 · 2026-08-27

UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City

Tianjie Ju, Zheng Wu, Yueqing Sun, Yuhan Cui, Bobo Li, Shengqiong Wu, Pengzhou Cheng, Haodong Zhao, Zongru Wu, Xinbei Ma, Doris Zhang, Kunling Li, Mong-Li Lee, Wynne Hsu, Hao Fei, Qi Gu, Gongshen Liu, Zhuosheng Zhang
具身AI空间智能MLLM Agent
💬 MLLM agent 认得出街景,但走着走着就迷路了——在 1:1 还原的香港城市沙盒里测试空间智能体的「组合失效」。

🎯 背景

多模态大模型能读懂单张街景图,但城市级的空间智能真正考验的是:局部感知在 agent 开始移动之后还有没有用。这个问题此前缺乏可测试的、物理约束真实的大规模城市环境。

🔬 方法

UrbanGround 基于香港全域三维地理数据构建首个物理约束真实城市沙盒,支持第一人称视角的闭环交互和地图导航;按三个递进问题设计实验:agent 主动观察后能否很好地为局部场景定位并回答空间问题?这种定位能力能否支撑距离更远、目标更模糊的导航?行为在路线可用性变化和行人干扰下是否还能保持?

📊 结果与意义

当代 MLLM agent 在视觉识别和短程空间推理上表现出有用的原子能力,但方向感和行人感知移动始终不可靠;核心失效出现在长程探索中——局部能力无法组合成持续的目标导向行为,误差不断累积且缺乏有效纠正机制,这对任何需要长时程空间决策的 embodied agent 都是警示。
▶ 原文摘要 Abstract
Multimodal large language models (MLLMs) can interpret a street view, but urban agency depends on whether such local evidence remains useful after the agent starts to move. In this paper, we investigate how far current MLLM agents can turn local urban perception into reliable action in a complicated real-scale city. We propose UrbanGround, the first sandbox to make this question testable in a physically constrained replica of Hong Kong built from territory-wide 3D geospatial data. UrbanGround supports closed-loop interaction from a first-person view and provides an interactive map for navigation. Agents can directly enter the 3D city and explore from a first-person view. Our analysis follows the growth of the spatial problem through three research questions. We first test whether an agent can ground a local scene well enough to answer spatial questions after active observation. Then we ask whether that grounding supports navigation as destinations become farther away and less explicit. Finally, we examine whether the resulting behavior survives changes in route availability and pedestrian motion. Contemporary MLLM agents usually show useful atomic abilities in visual recognition and short-range spatial reasoning, while orientation and pedestrian-aware movement remain unreliable. Their central failure emerges over extended exploration, where local abilities do not compose into sustained goal-directed behavior and errors accumulate without effective correction. We hope UrbanGround will support broader study of how far current MLLM agents can explore reliably in complex, open-ended urban environments.
来源:arXiv:2608.27456 · 精读基于摘要与 arXiv HTML/abs 页信息生成,未解析 PDF 全文