← 首页|学术|StateMemBench: Benchmarking Memory State Tracking in LLM Agents
cs.AI · cs.CL · 2608.19652 · 2026-08-20

Can Agent Memory Systems Track Evolving State?

Fan, Xinyi; Liu, Miri; Yang, Ruozhen; Ouyang, Siru; Han, Jiawei
一句话:记忆要跟上演化的状态——新基准专测记忆系统的动态更新能力

问题

LLM agent 记忆系统在长交互中面临状态演化问题:约束变化、决策被覆盖,但现有基准测的是静态回忆

方法

新基准:构造含状态演化的多轮交互场景,测量 agent 记忆系统能否反映最新状态而非历史累加

结果

揭示主流记忆系统在状态跟踪上的系统性缺陷;提出评测记忆「新鲜度」的新维度

与研究方向的关联

全双工 agent 中对话状态实时变化,需要记忆系统持续追踪而非静态积累

原文摘要

As LLM-based agents are deployed for longer and higher-stakes tasks, their memory systems continue to have crucial gaps. While existing memory benchmarks focus largely on recall-shaped tasks, we argue an effective memory system must track the evolving state of the world; as facts, constraints, and decisions are revised over a long interaction, answers must reflect the current state and not a superseded one. We define this capability as state tracking and instantiate it in StateMemBench, a benchm
MemoryState TrackingBenchmarkLong-Context
ArXiv 2026-08-22 日报精读 · 返回简报