← 首页|学术|Thinkingbox: Stateful Business Workflow Benchmark for LLM Agents
cs.CL · cs.DB · 2608.19741 · 2026-08-20

One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows

Li, Zhuochun; Ko, Youngmin; Keramati, Ali; Ferri, Nicola; Pelaez, Susana Palmaz Lopez
一句话:Pass@1 不够——Thinkingbox 测量 agent 在有状态业务流程中的真实可靠性

问题

现有 agent 基准测的是单次成功或代码/网页导航,忽略了企业工作流中多轮政策合规和依赖协调

方法

Thinkingbox:可执行 sandbox,业务工作流场景,测量多轮信息收集 + 策略遵守 + 依赖协调的综合可靠性

结果

主流 agent 在单步成功率上高估了实际工作流可靠性;多轮状态管理是主要瓶颈

与研究方向的关联

现实 agent 部署需要在长对话中维护工作流状态,与 duplex agent 的多轮上下文管理直接对应

原文摘要

Recent agent benchmarks increasingly ground evaluation in executable environments, from code repair to web navigation, app APIs, and function calling. Yet completing consequential work beyond code requires more than producing a plausible response or valid tool call: agents must gather missing information over multiple turns, follow domain policies, coordinate dependent tools, and realize the correct persistent state transition without collateral effects. In this paper, we introduce Thinkingbox,
BenchmarkBusiness WorkflowReliabilityMulti-Turn
ArXiv 2026-08-22 日报精读 · 返回简报