← 首页|学术|Hear2Act: Prosody-Driven Task Decision Benchmark
cs.CL · 2608.19515 · 2026-08-20

Hear2Act: Benchmarking When Prosody Should Change What an Assistant Does

Liu, Xinyi; Nayyeri, Hooshang; Hakkani-Tur, Dilek; Yilmaz, Emine; Kim, JK
一句话:韵律应改变行动,现有音频 agent 做不到——Hear2Act 基准量化这一差距

问题

现有基准把韵律感知、回复适当性、任务对话三部分割裂评测,无法检验韵律信息是否真正影响任务执行路径

方法

Hear2Act 基准:控制变量实验,同一语言内容配不同韵律,测量系统响应差异;覆盖多类任务对话场景

结果

量化了主流音频-LLM 在韵律驱动决策上的系统性缺失;为音频 agent 能力评测提供了新维度

与研究方向的关联

对话 agent 应根据用户语气调整行为,与全双工交互中的多模态信号处理直接相关

原文摘要

Prosodic cues can convey task-relevant information that alters the trajectory and outcome of a task-oriented dialogue, even when the words themselves remain unchanged. Yet existing benchmarks typically evaluate prosodic perception, response appropriateness, and task-oriented dialogue in isolation, making it difficult to test whether prosodic evidence changes downstream decisions. We introduce Hear2Act, a unified evaluation protocol for text and spoken assistants with 480 persona-grounded scenari
ProsodyDialogueAudio-LLMBenchmark
ArXiv 2026-08-22 日报精读 · 返回简报