LIBERO 40任务×4 WAM配对闭环干预:掩蔽/重赋future-cache值改变执行并降低成功率,但固定final-clean K/V缓存几乎保持未修改执行——分离「缓存消费」与「缓存生产」;RIFT = rollout-free imagination via future tokens。
章节结构(全文标题提取): 1 Introduction 2 Related Work 3 What does the action expert read? 4 Rift: One-pass Future-Token Imagination 5 Experiments 6 Conclusion References Appendix A Training implementation details 3.1 The channel we edit 3.2 Scoring each edit 3.3 Intervention set 3.4 Finding 1: WAM action experts use future values at their assigned positions 3.5 Finding 2: one final-clean K/V cache nearly preserves execution 4.1 From findings to our design 4.2 Writing the cache in one pass 4.3 Training 5.1 Setup 5.2 Quantitative results 5.3 Ablations 5.4 Qualitative results
World action models (WAMs) condition robot actions on predicted futures, but iterative video rollout increases deployment latency. We ask whether action generation requires the evolving rollout trajectory or only its future representation. Across four WAMs on all 40 LIBERO tasks, paired closed-loop interventions show that masking or reassigning future-cache values changes execution and reduces success, indicating sensitivity to future values and their assigned positions. For Joint and Cosmos-2, however, replaying one fixed final-clean key/value (K/V) cache nearly preserves unmodified execution, with $1.7$ to $1.9$~cm end-effector average displacement error and $97.9\%$ to $98.2\%$ success. This separates cache consumption from production: these models can reuse a fixed cache but still require iterative rollout to construct it. We therefore propose RIFT (\emph{Rollout-free Imagination via Future Tokens}), which uses learned anticipation tokens to construct a complete future K/V cache in one backbone pass while retaining the original future-read interface. On LIBERO, RIFT achieves $98.8\%$ success, close to rollout-based Joint, IDM, and LingBot-VA at $98.4\%$ to $98.6\%$, while reducing action-chunk latency by $68.2\%$ to $89.1\%$. On RoboTwin~2.0, RIFT reaches $92.9/92.6\%$ on clean/randomized scenes, the highest observed among the evaluated methods. These results support rollout-free future conditioning without iterative video generation at deployment.