← 首页|学术|AtlasVLA: Persistent World-Ego State Modeling for Vision-Language-Action Models
cs.RO,cs.CV · 2608.06729 · 2026-08-07

AtlasVLA: Persistent World-Ego State Modeling for Vision-Language-Action Models

Guiyu Zhao,Longteng Guo,Yanghong Mei,Zilin Zhu,Yu Zhang,Bin Cao,Mingming Yu,Xingjian He,Jie Jiang,Jing Liu
TL;DR:单腕相机下的 VLA 天然反应式,物体离开视野就"感知遗忘"。AtlasVLA 用持久世界状态记忆 + 自我工作状态记忆,把反应式操作变成带状态追踪的长程操作。

🎯 问题

VLA 推动了具身 AI,但其根本上的反应式范式在部分可观测、长程任务中严重受限。单腕相机下,物体离开视野即发生感知遗忘,多步执行中发生时间任务进度遗忘。

🔬 方法

AtlasVLA:持久世界状态记忆(Persistent World State Memory)+ 自我工作状态记忆(Ego-Working State Memory)+ 世界-自我引导的动作生成(World-Ego-Guided Action Generation),把反应式操控转为显式状态追踪。
章节结构(全文标题提取):
1 Introduction
2 Related Work
3 Method
3.1 Overview of AtlasVLA
3.2 Persistent World State Memory
3.3 Ego-Working State Memory
3.4 World-Ego-Guided Action Generation
4 Experiments
4.1 Experimental Setups
4.2 Simulation Evaluation on LIBERO

📊 结果

LIBERO / RLBench 仿真与真实环境评测均显著提升长程任务成功率,消融证明两种记忆各自必要。

📝 原文摘要

▶ 原文摘要 Abstract
While Vision-Language-Action (VLA) models have advanced embodied AI, their fundamentally reactive paradigm severely limits performance in partially observable and long-horizon tasks. When restricted to a single wrist-mounted camera, they inevitably suffer from perception forgetting as objects exit the field of view, and temporal task-progress forgetting} during multi-step execution. To overcome these bottlenecks, we propose AtlasVLA, a novel framework that transitions from direct reactive manipulation to proactive reasoning through a persistent world-ego state. AtlasVLA features a dual-memory architecture: a 4D Persistent World State Memory that lifts transient 2D observations into a globally updated, voxel-hashed spatial state to resolve visual blind spots, and an Ego-Working State Memory that tracks historical ego state and task progress. By conditioning a diffusion transformer (DiT) on this joint World-Ego state, AtlasVLA enables robust spatial reasoning. Extensive evaluations across LIBERO, RLBench, and real-world benchmarks demonstrate that AtlasVLA achieves state-of-the-art performance using solely a wrist camera. Remarkably, it decisively outperforms multi-view baselines, yielding absolute success rate improvements of 9.4% on LIBERO-Long and 17.5% in real-world long-horizon tasks.
Deep Read · 2026-08-11高松灯 / Agent 日报
EmbodiedVLAMemoryLong-Horizon