← 首页|学术|One Policy, Many Embodiments: Unified Camera-Centric Action Geometry Pre-training for Heterogeneous Embodied Manipulation
cs.RO · 2608.26058 · 2026-08-26

One Policy, Many Embodiments: Unified Camera-Centric Action Geometry Pre-training for Heterogeneous Embodied Manipulation

Xiaomi Embodied Intelligence Team, University of Macau, Shaoqing Xu, Fang Li et al.
VLA跨本体泛化动作表征
💬 异构机器人数据(不同本体、相机配置、动作空间)一直是scaling通用VLA策略的瓶颈,UCAG-P把动作统一表示为相机可观测的锚点运动,将机械臂、人形机器人、人手都视为同一几何动作schema的不同实例,单一checkpoint在LIBERO上达到98.3%。

🎯 背景

通用视觉-语言-动作策略的scaling受具身数据固有异质性的严重制约——不同机器人本体、相机配置和底层动作空间跨度极大。现有范式通常靠显式动作重定向、人到机器人视频合成或数据集特定适配分支来处理这种不匹配,从根本上阻碍了统一策略的联合学习。

🔬 方法

UCAG-P是一种以相机为中心的统一动作formulation,将异构具身数据集在结构上对齐进共享的几何动作空间。它不把机器人特定命令当作共享策略目标,而是通过图像和相机坐标系中相机可观测的锚点运动来表示操作,将机械臂、人形机器人和人手都视为同一通用动作schema的不同具身实例。一个几何条件的动作转换器将预测的运动与目标具身运动学结合,生成可执行控制;由此产生的解耦架构使共享VLA策略能学习可迁移的操作几何,同时保留具身特定的可控性。

📊 结果

UCAG-P在4.03千小时机器人与仿真数据、2.34千小时人类演示数据上训练,单一checkpoint在LIBERO上达到98.3%,在RoboTwin Easy/Hard上分别达到88.7%和89.2%,在LIBERO-Plus上零样本达到82.0%,在RoboCasa GR-1上达到62.0%,且均未做benchmark特定微调。

原文摘要

▶ Abstract
Scaling generalist vision-language-action (VLA) policies is severely bottlenecked by the inherent heterogeneity of embodied data, which spans diverse robot morphologies, camera configurations, and low-level action spaces. Existing paradigms typically address this mismatch through explicit action retargeting, human-to-robot video synthesis, or dataset-specific adaptation branches, fundamentally hindering the joint learning of a unified policy. We introduce UCAG-P, a camera-centric unified action formulation that structurally aligns heterogeneous embodied datasets into a shared geometric action space. Rather than treating robot-specific commands as the shared policy target, UCAG-P represents manipulation through camera-observable anchor motion in image and camera-frame coordinates, treating robot arms, humanoids, and human hands as different embodiments of a common action schema. A geometry-conditioned action translator combines predicted motion with target-embodiment kinematics to produce executable controls. The resulting decoupled architecture allows a shared VLA policy to learn transferable manipulation geometry while retaining embodiment-specific controllability. UCAG-P is trained on 4.03K hours of robot and simulation data and 2.34K hours of human demonstrations. A single checkpoint reaches 98.3% on LIBERO, 88.7% and 89.2% on RoboTwin Easy and Hard, 82.0% zero-shot on LIBERO-Plus, and 62.0% on RoboCasa GR-1, without benchmark-specific fine-tuning.
来源: arXiv:2608.26058 · 精读由高松灯生成,基于摘要与 arXiv 页面信息