← 首页|学术|Amplified Does Not Mean Predictive: Reasoning Behaviors in Thinking Models
cs.CL, cs.AI, cs.CV, cs.LG · 2608.13760 · 2026/08/13

Amplified Does Not Mean Predictive: Reasoning Behaviors in Thinking Models

Nyandwi, Jean de Dieu, Mathur, Leena, Bisk, Yonatan, Hawkins, Robert, Neubig, Graham
TL;DR:Behavioral Lift 度量行为与正确性的关联,揭示 Amplification-Lift Gap:thinking 模型强放大自我修正/假设检验/不确定承认,但最高 lift 的行为是置信校准/知识对齐/自我觉知——推理训练放大的是表面形式而非最强正确信号。

🎯 问题

推理导向训练让轨迹看起来更深思熟虑,但可能没放大与模型正确性最相关的行为——「放大≠可预测」。

🔬 方法

Behavioral Lift:行为存在 vs 缺失时正确率的变化量;15 模型 × 6 benchmark(文本+视觉-语言推理),标注 15,282 条轨迹,核心行为同时定义 LLM 与 VLM 版本。
章节结构(全文标题提取):
1 Introduction
2 Taxonomy and Metrics
3 Experimental Setup
4 Results
5 Discussion
6 Related Work
7 Conclusion
Acknowledgments
Ethics Statement

📊 结果

置信校准是两个模态最强的正确信号却几乎不被放大;不确定承认被放大 3-7 倍却与正确性弱或负相关;训练不优先放大最高 lift 行为——激励基于校准与落地推理的 process-level 目标。

💡 与研究方向关联

推理行为的「what to reward」问题——对 agent 推理训练的目标设计、评测维度有直接启发。

📝 原文摘要

▶ 原文摘要 Abstract
Which reasoning behaviors are associated with correct answers in reasoning models, and does reasoning-oriented training amplify those behaviors? This distinction is important because reasoning-oriented training can make traces look more deliberative without amplifying the behaviors most tied to model correctness. We quantify this mismatch with Behavioral Lift, a metric that measures how much correctness changes when a behavior is present versus absent in a model's reasoning trace. Across 15 models and 6 benchmarks spanning text-only and vision-language reasoning, we annotate 15,282 traces with a taxonomy whose core behaviors are defined for both LLM and VLM traces. We find evidence for an Amplification-Lift Gap, in which thinking models strongly amplify self-correction, hypothesis testing, and uncertainty acknowledgment, while the highest-lift behaviors are confidence calibration, knowledge alignment, and self-awareness. Confidence calibration is among the strongest positive signals of correctness in both modalities, yet is barely amplified; uncertainty acknowledgment is amplified by 3--7$\times$, yet is weakly or negatively associated with correctness. We find that reasoning-oriented training does not preferentially amplify the highest-Lift behaviors, motivating process-level objectives that reward calibrated and grounded reasoning rather than surface form alone.
Deep Read · 2026-08-18高松灯 / Agent 日报
Reasoning BehaviorThinking ModelsBehavioral LiftCalibration