← 首页|学术|ArXiv 日报 — 2026-08-28
学术 · 2026-08-28

ArXiv 日报

Agent 架构 · Memory · RL训练 · 具身智能
要点:今日从279篇候选标题中筛选29篇精读候选、最终精选16篇深度精读,分为4个主题。全双工/duplex agent 关键词扫描:对全部279篇候选标题逐一扫描,命中 0 篇(本轮无duplex相关新论文)。亮点包括用单一线性方向调节工具调用率的表征编辑方法、去中心化多智能体自发涌现社会分工的SwarmWorld实验,以及揭示参数化记忆"存储可行但检索失效"的知识图谱LoRA研究。知识库板块:AgentTeam-Shared-Knowledge 本地仓库已 pull 至最新(HEAD 8c336b2),checkpoint之后暂无新增 commit。
279
候选标题
29
精读候选
16
精选精读
0
Duplex命中
  1. 🧩 Agent 架构与协作 (5篇)
  2. 🧠 Agent Memory (4篇)
  3. ⚡ RL训练与推理效率 (4篇)
  4. 🤖 具身智能 / VLA (3篇)

🧩 Agent 架构与协作

Tunable Tool-Call Rates in LLM Agents via Representation Steering
cs.AI · 2608.25198 · Yuqi Chen, Vincent Siu, Yang Liu, Dawn Song et al.
用残差流中的单一线性方向就能把工具调用率从接近0%连续调节到90%以上,无需训练、不改prompt,还能泛化到未见过的工具,与duplex场景里"该不该触发某个实时动作"的决策机制高度相通。
▶ 原文摘要 Abstract
Deciding whether to call a tool is a core competence of an LLM agent, and a costly one to get wrong: needless calls add latency, accrue cost, and may trigger irreversible side effects, while missing calls leave the model confidently wrong on questions it could only answer through tool-calls. Models manage this balance poorly, both over-using and under-using tools. Existing methods such as post-training and prompt engineering are expensive and difficult to modify at inference time. We show that whether an instruction-tuned model calls a tool can be controlled by a single linear direction in its residual stream, extracted without any training from the model's own tool-use preference signal and turned into an inference-time intervention with no prompt change. Adding the direction with strength $\alpha$ moves the call rate monotonically from near $0\% $ to over $90\%$ while keeping calls well-formed. The steering works in both directions: dialing it down suppresses calls, and dialing it up induces new calls that land precisely on the questions the model cannot answer from its own knowledge. We also show that the direction generalizes to unseen tools with strength comparable to each tool's own direction and without favoring any specific tool choice. With live tool execution, a single sweep of the steering traces a cost/accuracy Pareto frontier and nearly doubles open-domain QA accuracy ($0.29 \! \rightarrow \! 0.56$); the same recipe transfers across a diverse range of models spanning dense, MoE, and multimodal architectures, without any training. Our code is publicly available at this https URL .
工具调用表征编辑推理时控制
SwarmWorld: Stigmergic technological evolution in societies of language-model agents
cs.AI, cond-mat.mtrl-sci, cs.CL · 2608.26081 · Subhadeep Pal, Fiona Y. Wang, Markus J. Buehler
让同质化LLM agent在共享环境中不设角色、不给脚本,观察其能否靠"物理痕迹"而非直接对话自发分化出社会分工并演化出可持续技术体系——这是对"多智能体是否需要显式协议"的一次极端实验。
▶ 原文摘要 Abstract
Collective intelligence can emerge when individuals coordinate through a shared environment, allowing local actions to accumulate into durable social organization. Language-model agents offer a new substrate for this process, yet most multi-agent systems rely on direct conversation, predefined roles, or centralized workflows. It remains unclear whether decentralized agents can build functional technologies and outperform independent search. Here, initially homogeneous LLM agents in SwarmWorld self-organize without assigned roles or recipes into evolving technological societies. Agents explore a spatial environment, process resources, test materials, construct persistent artifacts, and write executable controllers evaluated by a deterministic simulator under unseen disturbances after the agents are removed. SwarmWorld splits cognition from consequence: agents propose architectures and controllers within fixed action and material schemas, while the simulated world determines function. Shared societies develop broader, more resilient technological portfolios than a strong best-of-N isolated-search baseline, although isolated search remains competitive for the strongest artifact. Agents differentiate into exploration, construction, maintenance, and coordination behaviors, transitioning as the world matures. Technologies accumulate through collaborative construction, executable inheritance, and persistent agent-artifact networks, with most reuse beginning through physical observation rather than communication. Explicit cultural mechanisms amplify collaboration and organization, but functional benefits depend on outcome and timescale. Physical stigmergy alone supports capable societies, while interaction drives persistent technological ecologies rather than universally superior individual inventions.
多智能体群体智能涌现行为stigmergy
JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution
cs.CL, cs.LG · 2608.25593 · Guibin Zhang, Leo Lu, Fangzhou Xie, Kang Zhu et al.
把agent harness(记忆管理/规划策略/动作协议/工具编排)本身当作可训练、可即时生成的产物,训练出专门模型为任意基座LLM实时合成/修复/自我进化harness,证明"harness智能"是独立于模型规模的可训练维度。
▶ 原文摘要 Abstract
Agent capability is not determined by the model alone. The agent harness, encompassing memory management, planning strategy, action protocol, and tool/skill orchestration, can dominate the contribution of the underlying foundation model. Yet harness design remains manual, task-specific, and fundamentally unscalable. We present JIT-Agent, a harness intelligence model trained to synthesize task-adaptive agent harnesses on the fly for arbitrary off-the-shelf agentic LLMs. We formalize the agent harness as a composable, machine-generatable artifact governed by a fixed four-module protocol, and train JIT-Agent to customize harnesses for a given task at hand, repair harnesses for stable and reliable execution, and self-evolve by distilling performance signals from an expanding archive of prior harness configurations. Equipped with JIT-Agent as a harness helper, DeepSeek-V4-Flash surpasses GPT-5.6 on DeepSearchQA (+9.1) and OdysseyBench (+4.3), while the already strong GLM-5.2 gains up to +20.2 points. Across controlled evaluations, JIT-Agent-generated harnesses are performance-competitive with mature agent runtimes such as OpenCode and Claude Code and consistently improve multi-scale model families of DeepSeek V4, Mimo-V2.5, and Qwen3.6. To our knowledge, JIT-Agent is the first model purpose-built for just-in-time harness generation, establishing harness intelligence as a trainable, transferable, and compounding dimension of agent capability orthogonal to model scaling.
Agent架构Harness自我进化
TOPAS: Workflow-Aware Prefix-State Scheduling for Multi-Agent LLM Serving
cs.CL · 2608.25523 · Hongqiu Ni, Han Tian, Chi Zhang, Guopeng Li et al.
多智能体LLM服务中,给某agent保留长system-prompt的KV缓存能加速后续调用但挤占并发显存,TOPAS联合决定"缓存哪些前缀"与"调度哪些请求",把任务完成时间均值/p99最多降低约40%/49%,是支撑低延迟多智能体交互的基础设施拼图。
▶ 原文摘要 Abstract
Prefix caching introduces a fundamental tradeoff in multi-agent large language model (LLM) serving: retaining a long system-prompt key-value (KV) cache for an agent accelerates future calls, yet it reduces the GPU memory available for batching concurrent requests. In multi-stage workflows, existing schedulers tend to prioritize either immediate prefix locality or overall workflow progress. However, under a shared KV cache budget, optimizing either objective in isolation can prolong tasklevel job completion time (JCT) through downstream delays or frequent prefix replacement. To strike a balance, we here propose TOPAS, a Task-Oriented Prefix-Aware Scheduler that jointly decides which agent prefixes to keep in the cache and which requests to schedule for execution. TOPAS scores candidate post-decision states by trading off the expected reduction in each task's longest remaining service path against the near-term benefit of downstream prefix reuse, accounting for the costs of prefix movement and preemption. A task-level aging mechanism is also incorporated to prevent starvation. We implement TOPAS within the SGLang framework and assess its performance on three synthetic DAGs and two MetaGPT software-development workflows. Compared with the best performing baseline for each workload and metric, TOPAS reduces the mean/p99 JCT by up to 39.8%/49.4% on the synthetic workloads, while lowering mean JCT by 9.8% on MetaGPT-SOP and mean/p99 JCT by 22.0%/26.6% on MetaGPT-TL.
推理基础设施KVCache多智能体调度
Repair or Resample? Rethinking Failure Debugging in LLM Multi-Agent Systems
cs.AI, cs.SE · 2608.25920 · Zhongwen Luan, Xiaoyu Zhang, Ming Hu, Yue Yang et al.
现有多智能体系统"修复"方法到底是真的因果修复了失败,还是只是靠LLM采样随机性侥幸绕过?作者发现无引导重跑的失败复现率仅68%、修复率仅6.9%,而症状驱动的干预方法能把修复率提升191%——这是给所有构建多智能体系统的人的一记警钟。
▶ 原文摘要 Abstract
As large language model (LLM)-based multi-agent systems (MASs) are increasingly applied to long-horizon complex tasks, their reliability has emerged as the core bottleneck hindering their real-world deployment. Existing MAS debugging and repair methods typically rely on rerunning and resampling the entire execution trajectory. However, a fundamental question remains to be answered: do these methods causally repair MAS failures or merely stochastically repair by leveraging the randomness of LLM sampling? To evaluate the effectiveness of MAS repair methods, we introduce SymTrace, a controlled evaluation framework that records the MAS execution trajectory and establishes intervention anchors. During replay, it effectively reconstructs the execution before the anchor using recorded logs and only regenerates the downstream trajectory, thereby enabling the reliable reproduction of MAS failures. We further construct the dataset SymFail, comprising 536 human-annotated failure trajectories with graph-linked locations, categories, and trace evidence. Based on these foundations, we conduct a large-scale empirical study across three mainstream MAS frameworks. Our findings reveal that existing unguided rerun methods are highly unreliable, exhibiting low failure reproduction and repair rates (only 67.97% and 6.90%, respectively). Building upon these findings, we further explore the effectiveness of a symptom-driven intervention method, which successfully repairs 20.15% of the failed cases (a 191.89% improvement to state-of-the-art repair methods). This study aims to provide actionable insights for MAS debugging and repair research, paving the way for the robust deployment of multi-agent systems.
多智能体调试可靠性

🧠 Agent Memory

Learning What to Share and What to Personalize: Hierarchical Strategy Co-Evolution for Agent Memory
cs.AI, cs.CL · 2608.25329 · Yupeng Han, Shuochen Liu, Kai Zhang, Ze Liu et al.
记忆管理策略该"一刀切"还是"千人千面"?HiPS把记忆决策拆成全局共享原则与用户特异调整两层,并用跨层规则流动态校准边界,为长程/持续交互agent的记忆架构设计提供了可参考的在线权衡范式。
▶ 原文摘要 Abstract
Memory-augmented agents maintain compact user profiles throughout extended conversations, enabling personalized and consistent responses without the need to process the entire dialogue history. The quality of these user profiles relies on the underlying memory management strategy: at each step, the agent must determine what to retain, compress, or discard. However, existing methods typically employ a static, one-size-fits-all strategy established before training. In practice, the optimal memory decision is inherently user-specific and dynamically evolves alongside policy optimization. To address this, we propose \textbf{HiPS} (\textbf{Hi}erarchical \textbf{P}ersonalized \textbf{S}trategy), a framework that decouples memory management into a globally shared foundation and a user-specific adaptive tier. Specifically, HiPS employs \textbf{Universal Strategy} to extract shared principles from cross-persona trajectories, alongside \textbf{Persona Delta Distillation} to generate tailored rules for users whose behaviors diverge from general patterns. \textbf{Cross-Level Rule Flow} dynamically calibrates their boundary by promoting broadly validated personal rules and demoting contradicted global ones. The architecture establishes a co-evolution loop where a mechanism guarantees that all strategy refinements are anchored to task outcomes. Extensive experiments demonstrate consistent improvements over memory-augmented baselines.
Agent记忆个性化策略共演化
A Storage-Retrieval Gap in Parametric Knowledge Graph Memory
cs.LG, cs.CL, cs.IR · 2608.25489 · Martino M. L. Pulici, Cuong Xuan Chu, Evgeny Kharlamov, Volker Tresp
把知识图谱离线编译成每实体一个LoRA adapter、注入权重而非塞进context来查询,零查询时上下文开销;但发现存储的知识无法靠相似度检索——知识存储是局部的、不可迁移的,揭示了参数化记忆的一个根本性开放问题。
▶ 原文摘要 Abstract
Graph retrieval-augmented generation places retrieved subgraphs into the model's context window at query time, paying a recurring token cost and exposing source data on every call. We study an alternative: compiling a knowledge graph offline into a bank of LoRA adapters, one per entity, that serve as a parametric knowledge layer queried by injecting weights rather than text, at zero query-time context cost. On the MetaQA dataset, we find that subgraph-trained adapters encode context-free factual knowledge that generalizes to unseen questions: on single-valued relations the adapter gains $+0.243$ exact-match score over a base model that is nearly blind closed-book ($0.007$), and only the correct adapter recovers this knowledge (an oracle gap of $+0.283$ over the base model). However, the stored knowledge is not recoverable by similarity: given a query with no subgraph, embedding-based and weight-space geometry retrieval both perform at chance, because a semantically neighbouring entity's adapter does not contain the answer - knowledge is stored locally and does not transfer. Weight geometry correlates with subgraph semantics ($\rho = +0.329$) but not with functional retrievability. We quantify the byte and context-token costs against graph retrieval-augmented generation and discuss deployment implications. Our results establish that parametric knowledge graph memory is feasible for storing knowledge, and identify selecting and composing the right adapters by a mechanism other than semantic similarity as the central open problem - motivating a learned, query-conditioned composition mechanism.
参数化记忆知识图谱LoRA
AWM: Answerable Working Memory for Long-Document VQA Agents
cs.CL · 2608.25618 · Dongzhuoran Zhou, Yuqicheng Zhu, Yule Liu, Zhen Yang et al.
agent可能翻对了页、答对了问题,但留下的工作记忆本身太笼统或不完整,脱离原始页面就无法支撑答案——即使给了黄金证据页,42.5%的正确答案其实无法只靠终态工作记忆回答;把这个"记忆可回答性"信号纳入GRPO奖励后准确率提升8-12个点。
▶ 原文摘要 Abstract
Long-document visual question answering increasingly relies on VLM agents that retrieve candidate pages, inspect page images, write findings to working memory, and synthesize answers. Working memory should carry answer-supporting evidence across page inspections for later grounded answering, yet existing evaluation mainly checks final-answer correctness and evidence-page access. This creates a memory-quality blind spot: an agent may reach the right page and answer correctly while leaving behind memory too generic or incomplete to support answering once page context is removed. We introduce \emph{memory-only answerability}, a diagnostic that asks whether a reader can answer from the question and terminal working memory alone. Building on this diagnostic, \emph{Answerable Working Memory} (AWM) treats terminal working memory as an answerable evidence artifact, and AWM-GRPO incorporates this signal into the GRPO reward while preserving final-answer priority. Under GRPO, this reward assigns higher advantages to answer-correct trajectories whose terminal working memory remains answerable. On \textsc{MMLongBench-Doc}, even when gold evidence pages are provided, 42.5\% of correct answers still cannot be answered from terminal working memory alone. AWM-GRPO improves final-answer accuracy over the RAG baseline by 8.1 and 11.9 points on \textsc{MMLongBench-Doc} and \textsc{LongDocURL} and reduces the memory-missing-correct rate by 2.7 points over answer-only GRPO.
Agent记忆工作记忆VQA
Reconstructing the Right Episode: Evaluating Interleaved Conversational Memory Beyond Long Context
cs.CL, cs.AI · 2608.25655 · Zhexi Feng, Ruiyi Zhang, Yongbo Yang, Pengtao Xie
现实对话助手往往是不分session、话题混杂的长线程,系统必须自行判断"哪段更早的对话内容决定了当前任务决策是否成立"——这种扁平交错线程中的情节完整性问题,和duplex场景下持续、无边界的实时交互记忆需求几乎同构。
▶ 原文摘要 Abstract
Conversations with chat assistants increasingly span many topics in a single long-running thread, challenging memory systems. Existing long-context and memory benchmarks often expose session or topic boundaries, or probe direct personal-memory questions. These settings understate a harder assistant-memory regime: a flat mixed-topic thread where the system must infer which earlier episode makes a later task decision valid. We introduce SCALE-QA, a constraint-grounded task QA benchmark for flat unsegmented threads targeting episode integrity failure. The dataset contains 3,000 audited questions across 10 domains, uses deterministic four-way multiple-choice grading, and includes a deterministic runtime builder; experiments use all 3,000 questions through 128k and a stratified 400-question diagnostic at 1M. SCALE-QA questions are ordinary task-oriented requests whose correct answer depends on causally related evidence introduced earlier in the conversation. We also propose Temporal-Semantic Interleaved Memory Reconstruction (TSIM), which segments the turn stream into coherent episodes and indexes them through a hierarchical multi-view memory stack with deterministic episode-level summary and cluster-routing views. Experiments show that SCALE-QA challenges strong RAG baselines and long-context LLMs alike; across three open-source and proprietary LLM backends, TSIM achieves the highest accuracy in every backend setting, gaining 5.6-17.6 accuracy points over the strongest corresponding baseline.
对话记忆长上下文情节分割

⚡ RL训练与推理效率

Reflection Steering: Disentangling Reflection from Reasoning in Activation Space for Token-Efficient Inference
cs.LG, cs.CL · 2608.25542 · Jiarui Hu, Zhiyuan Wen, Xiaoyun Liu, Jiaxing Shen et al.
大推理模型常做无意义的"自我复查"浪费token,本文通过对比反思/非反思隐藏状态并与通用推理方向正交化,分离出纯粹的"反思方向",平均减少16.9%推理token,对延迟敏感的实时交互系统是即插即用的手段。
▶ 原文摘要 Abstract
Large reasoning models often produce reasoning traces with verification, revision, and backtracking. When reflection merely re-checks established results, it wastes reasoning tokens and increases latency. Most existing reflection steering methods add a label-derived mean-difference direction across preset layers, but its entanglement with reasoning and length signals destabilizes the accuracy-efficiency trade-off. In this paper, we propose Reflection Steering, a training-free framework for controlling reflection-associated computation within LLMs by disentangling reflection-related activations from general reasoning. Specifically, we contrast reflective and non-reflective hidden states at each LLM layer, denoise the resulting reflection directions with PCA, and orthogonalize them against general-reasoning directions. To limit downstream amplification from early-layer interventions, we calibrate each layer across multiple intervention strengths on a small set, retain only stable layers, and apply bounded projection removal to their residual-stream activations. We conduct extensive experiments across two public benchmarks and three open-weight LLMs against state-of-the-art activation-steering baselines. Results show that Reflection Steering reduces reasoning tokens by 16.9% on average across six matched settings. Besides, our method further introduces a bounded reflection intervention-strength parameter $\alpha$, enabling deployment-time adjustment to balance token savings, accuracy, and generation stability.
推理效率激活空间编辑延迟优化
AsymSpec: Context-Asymmetric Speculative Decoding for Agentic LLMs
cs.AI, cs.CL · 2608.26004 · Sheng Liang, Yongyue Zhang, Nathanael Brian, Hang Lv et al.
投机解码通常要求草稿模型和验证模型看到相同上下文,AsymSpec让轻量草稿模型读完整输入、大验证模型只读压缩视图,在多个agentic benchmark上达到约90%全上下文精度,吞吐提升1.3-1.7倍、算力降到0.2-0.3倍。
▶ 原文摘要 Abstract
Agentic LLM pipelines face escalating inference costs as context accumulates across retrieval, tool use, and multi-turn interactions. To control latency, deployments routinely compress inputs, but this degrades task accuracy. Speculative decoding (SD) accelerates generation losslessly, yet it assumes the drafter and verifier share an identical context, preventing SD from resolving the accuracy-overhead trade-off. We propose AsymSpec, an asymmetric speculative decoding framework that breaks this symmetry: a lightweight drafter reads the full input while the large verifier operates on the compressed view. The drafter steers the verifier via a contrastive $\delta$-fusion of logits, modulated by a divergence-aware acceptance gate that preserves verification stability and high draft acceptance rates. Evaluated across four agentic capabilities and two end-to-end agent benchmarks, AsymSpec reaches $\approx 90\%$ of full-context accuracy on average, delivering $1.3$--$1.7\times$ throughput speedups at $0.2$--$0.3\times$ the compute cost on isolated text capabilities. These results show that asymmetric context access yields substantial gains precisely when compression discards critical reasoning signals.
投机解码推理加速上下文压缩
Spectral Allocation: Why Muon Outperforms Adam, and How to Improve Muon
cs.LG · 2608.25990 · Xiaodong Wu, Wenyi Yu, Chao Zhang, Philip Woodland
对训练轨迹上真实checkpoint做频谱探测,发现动量缓冲的奇异值方向存在稳定的各向异性——"头部"方向在稳定性边缘只能用小步长、"体部"方向容忍度更高,这统一解释了Muon为何优于Adam;据此提出的SAMuon比调优过的Muon节省13-24%训练token。
▶ 原文摘要 Abstract
Orthogonal optimisers such as Muon can substantially accelerate large language model pretraining relative to Adam, yet the mechanism remains incompletely understood. We investigate this through an out-of-sample spectral probing analysis of Transformer loss landscapes. At checkpoints along real training trajectories, we decompose each momentum buffer into its singular directions and estimate the loss-optimal step size along each direction on held-out data. The resulting spectral profile is anisotropic yet stable across batches and training stages, and consistent across the optimisers and model scales: a volatile head operating at the Edge-of-Stability supports a much smaller step size than the tolerant bulk, which permits substantially larger steps. This profile provides a unified spectral allocation account of why Muon outperforms Adam, which outperforms SGD. It also exposes a limitation of Muon's uniform scaling: it still underutilises the bulk. Guided by this finding, we introduce Spectral-Aware Muon (SAMuon), which holds the head at the Muon scale and amplifies the bulk using a static spectral prior. We provide two variants: the complete SAMuon follows the measured profile using a low-rank randomised SVD and the simplified SAMuon-lite uses a two-level approximation via rank-one power iteration. Neither method adds persistent optimiser state or notable extra FLOPs beyond Muon at scale, and the idealised exact-whitening versions of both retain Muon's asymptotic convergence rate under standard assumptions. Across "modded-nanogpt" models from 124M to 1B parameters, both variants outperform tuned AdamW and Muon (Scion implementation) baselines in all evaluated model-scale and batch-size configurations. SAMuon requires 13.3% to 24.0% fewer training tokens to reach the same validation loss as Muon, while SAMuon-lite retains most of this gain with near-zero wall-clock overhead.
优化器Muon训练效率
One Symptom, Three Levers: A Critical Review of On-Policy Self-Distillation
cs.LG, cs.AI, cs.CL · 2608.25936 · Justin Robert, Raheel Qader
On-Policy Self-Distillation用模型自己当教师省掉独立教师模型成本,但同一种不对称性带来了"坍缩"——推理路径逐渐收窄,这篇综述把坍缩归纳成信号位置/教师信息/教师动态三个可调杠杆,为跨论文各说各话的现象提供统一词汇表。
▶ 原文摘要 Abstract
On-policy distillation trains a language model on its own generations while a teacher scores them token by token. It combines the dense supervision of imitation learning with the on-policy sampling of reinforcement learning. But it requires a second, larger model to act as teacher. On-Policy Self-Distillation (OPSD) removes that cost. The teacher is the model itself, conditioned on privileged information the student will not have at test time, such as a reference solution, a plan, or environment feedback. The teacher is no stronger than the student, only better informed. Early results were promising, with accuracy comparable to reinforcement learning at a fraction of the generated tokens. But the same asymmetry that produces the signal also biases it. One failure mode now dominates the field: collapse, the progressive narrowing of the set of reasoning paths the model can produce. Collapse is not specific to OPSD, though privileged information aggravates it. This review treats collapse as a symptom governed by three levers: (i) where the signal is applied, that is, how tokens are weighted; (ii) what the teacher is shown, that is, the nature of the privileged information; and (iii) when the signal changes, that is, the teacher's dynamics and the decay of guidance. We restrict our scope to mathematical reasoning, where the method originated and where its failure modes are best documented. We report no new experiments. The contribution is structural: a shared vocabulary for phenomena named differently across papers, and a clear line between what is settled and what is still disputed.
自蒸馏强化学习综述

🤖 具身智能 / VLA

$R^3$: Training Robots to Reason in Natural Language via Reinforcement Learning
cs.RO, cs.AI, cs.CL, cs.LG · 2608.26053 · Lehong Wu, Yuxiao Qu, Zheyuan Hu, Ivan Zhang et al.
让VLM先在专家推理轨迹上中期训练建立推理风格,再用基于规则的单步RL在离线动作数据上强化,把自由形式语言推理直接训练成指导底层操作策略的测试时计算机制,为"语言推理能否真正提升机器人长程操作"给出较为正面的答案。
▶ 原文摘要 Abstract
Reasoning in language allows foundation models to spend more test-time compute on hard problems, such as those requiring decomposition, constraint tracking, and prediction of future consequences. Whether this mechanism can improve robotic manipulation remains unclear, where long-horizon tasks require tracking partial progress, reasoning about object relations, recovering from mistakes, and steering noisy low-level policies. In this paper, we study whether VLMs can be trained to reason directly in natural language to guide low-level manipulation policies. We introduce $R^3$, a simple post-training recipe that turns off-the-shelf VLMs into robotic reasoners: it first mid-trains a VLM on expert-generated reasoning traces to initialize the desired reasoning style, then improves the reasoner with single-step rubric-based RL from offline action data. Unlike prior robotic reasoning methods that mostly use structured traces as auxiliary supervision, $R^3$ trains free-form language reasoning to produce test-time guidance for action. We instantiate $R^3$ on Language Table and simulated bimanual grocery packing, two controlled testbeds for studying robotic reasoning and long-horizon manipulation. $R^3$ improves exploration and generalization across unseen tasks and significantly outperforms instruction-only imitation learning baselines on both benchmarks. Our analyses suggest that free-form language reasoning can function as a test-time compute mechanism for steering low-level policies. Our project page is available at this https URL .
具身智能VLA强化学习语言推理
StreamPI: Streaming Multimodal Temporal Modeling for Vision-Language-Action Models
cs.CV · 2608.26067 · Zhe Liu, Jinghua Hou, Yuxiang Lu, Zhenya Yang et al.
现有SOTA VLA(如pi0.5)多是单帧范式、难保留历史观测,StreamPI用指令锚定的时间建模、零新增参数给单帧VLA装上时序推理能力,并用随机间隔流式训练提升对真实机器人异步部署帧间隔扰动的鲁棒性,其"流式+低延迟+异步部署"设计目标与duplex交互高度同构。
▶ 原文摘要 Abstract
Vision-Language-Action (VLA) models have demonstrated effectiveness in robot manipulation, yet state-of-the-art models such as pi0.5 operate under a single-frame paradigm, limiting their ability to retain past observations and develop precise spatial perception. In this paper, we propose StreamPI, a streaming multimodal temporal modeling framework that equips single-frame VLA with temporal reasoning capability without introducing any additional parameters. One core design is instruction-anchored temporal modeling. It treats each (visual observation, language instruction) pair as an atomic temporal unit: bidirectional attention within each pair enables cross-modal fusion, while causal attention across pairs preserves autoregressive streaming inference. This ensures the language instruction serves as a persistent semantic anchor throughout task execution. To bridge the gap between synchronous training and asynchronous real-robot deployment, we introduce a andom-interval streaming training strategy: a proper inter-frame interval (e.g., every 3 frames) enables faster and smoother action execution. Beyond this, randomizing the interval further improves robustness to frame-timing perturbations, supporting asynchronous deployment in practice. Furthermore, by leveraging the length extrapolation capability of the LLM backbone, StreamPI seamlessly inherits pretrained single-frame weights and supports flexible single-frame and multi-frame inference. Experiments on real-robot tasks spanning memory-dependent and precise perception scenarios, as well as the simulation benchmark LIBERO, demonstrate that StreamPI outperforms pi0.5 across diverse tasks.
VLA流式推理时序建模
One Policy, Many Embodiments: Unified Camera-Centric Action Geometry Pre-training for Heterogeneous Embodied Manipulation
cs.RO · 2608.26058 · Xiaomi Embodied Intelligence Team, University of Macau, Shaoqing Xu, Fang Li et al.
异构机器人数据(不同本体、相机配置、动作空间)一直是scaling通用VLA策略的瓶颈,UCAG-P把动作统一表示为相机可观测的锚点运动,将机械臂、人形机器人、人手都视为同一几何动作schema的不同实例,单一checkpoint在LIBERO上达到98.3%。
▶ 原文摘要 Abstract
Scaling generalist vision-language-action (VLA) policies is severely bottlenecked by the inherent heterogeneity of embodied data, which spans diverse robot morphologies, camera configurations, and low-level action spaces. Existing paradigms typically address this mismatch through explicit action retargeting, human-to-robot video synthesis, or dataset-specific adaptation branches, fundamentally hindering the joint learning of a unified policy. We introduce UCAG-P, a camera-centric unified action formulation that structurally aligns heterogeneous embodied datasets into a shared geometric action space. Rather than treating robot-specific commands as the shared policy target, UCAG-P represents manipulation through camera-observable anchor motion in image and camera-frame coordinates, treating robot arms, humanoids, and human hands as different embodiments of a common action schema. A geometry-conditioned action translator combines predicted motion with target-embodiment kinematics to produce executable controls. The resulting decoupled architecture allows a shared VLA policy to learn transferable manipulation geometry while retaining embodiment-specific controllability. UCAG-P is trained on 4.03K hours of robot and simulation data and 2.34K hours of human demonstrations. A single checkpoint reaches 98.3% on LIBERO, 88.7% and 89.2% on RoboTwin Easy and Hard, 82.0% zero-shot on LIBERO-Plus, and 62.0% on RoboCasa GR-1, without benchmark-specific fine-tuning.
VLA跨本体泛化动作表征

📚 知识库更新

AgentTeam-Shared-Knowledge 仓库已成功 git pull --ff-only,本地 HEAD 保持在 8c336b2(2026-08-26 提交),checkpoint 与当前 HEAD 一致,本轮 checkpoint..HEAD 区间内暂无新增 commit,无需展示增量内容。
来源: arXiv (cs.AI/cs.LG/cs.MA/cs.RO/cs.CL/cs.SE/cs.CV/eess.AS) · 简报由高松灯生成