← 首页|学术|不是铁板一块:中国前沿 LLM agent 的合作均衡差异
cs.MA · 2608.10262 · 2026/08/10

Not a Monolith: Lab-Level Divergence in the Cooperative Equilibria of Chinese Frontier LLM Agents

Bolívar, Francisco León Zúñiga
TL;DR:中国前沿 LLM agent 不是铁板一块:固定转换器(GPT-5.4 Mini)消除「策略 vs 编码能力」混淆后,四家实验室(DeepSeek V4 Pro / Qwen3-Max / Kimi K2.5 / GLM-5.1)在演化囚徒困境中激进均衡比例差异显著(1%-9%),生态内差异甚至大于东西方均值差。

🎯 问题

西方前沿 LLM agent 表现出的合作偏置是否适用于不同的对齐谱系?中国模型应被当作单一集团还是独立实验室?

🔬 方法

演化迭代囚徒困境 + 固定转换器设计(所有模型用同一 GPT-5.4 Mini 把自然语言策略转成代码,纯比较生成能力),全对全锦标赛 + Moran process n=500,3 种 prompt 风格 × 4 种群制度,预注册假设。
章节结构(全文标题提取):
1. Introduction
2. Related Work
3. Method
4. Results
5. Discussion
6. Conclusion
Data and Code Availability
Acknowledgements
AI Use Disclosure
3.1. Strategy Generation and the Conversion Confound · 3.2. Models · 3.3. IPD Tournament · 3.4. Attitude-Agents · 3.5. Moran Process · 3.6. Derived Metrics · 3.7. Pre-registered Hypotheses · 4.1. Strategy Validation: Cooperation Propensity

📊 结果

H6「非铁板」得到支持:P_A 从 1%(Qwen3-Max)到 9%(DeepSeek V4 Pro),4/6 配对比较通过 Holm-Bonferroni;生态内差异(8pp)大于中西方生态均值差(5.0% vs 5.0%)。

💡 与研究方向关联

多智能体系统研究关心模型间的合作/竞争均衡。该文的方法论亮点是固定转换器——剥离编码能力干扰,让跨模型比较纯粹反映生成倾向,这是同类研究的常见混淆源。

📝 原文摘要

▶ 原文摘要 Abstract
Does the cooperative bias documented for Western frontier LLM agents extend to a different alignment lineage, and should the Chinese models that embody it be treated as a single bloc or as distinct laboratories? We study four frontier-tier Chinese models - DeepSeek V4 Pro, Qwen3-Max, Kimi K2.5 and GLM-5.1 - in an evolutionary Iterated Prisoner's Dilemma, under a design that removes a confound present in prior work. Rather than letting each model convert its own natural-language strategies into code, which entangles strategic disposition with coding ability, we hold the converter fixed (GPT-5.4 Mini) across all labs, so every cross-lab comparison is a comparison of generation alone. We run the full protocol: all-play-all tournaments and a Moran process at n=500 runs per condition, across three prompt styles and four population regimes. Two pre-registered hypotheses are evaluated. H6 (not monolithic) is supported: the four labs differ significantly in aggressive-equilibrium proportion, P_A running from 1% for Qwen3-Max to 9% for DeepSeek V4 Pro, with four of six pairwise comparisons surviving Holm-Bonferroni. The spread across the four labs (P_A range 8pp) is larger than the difference between the Chinese and Western ecosystems' mean P_A (5.0% vs 5.0%): on this measure, within-ecosystem variation exceeds the East-West gap. H5 (cooperative-bias generality) is consistent but qualified: a cooperative plurality holds in 6 of 12 lab-prompt combinations against the 9 of 12 reported for Western models, a difference we do not treat as firm, since the count rests on Cooperative-Neutral near-ties and rises to 9/12 under an alternate converter in our pre-registered robustness check. The lab, not the ecosystem, is the unit at which cooperative disposition is set; treating "Chinese models" as a monolith is not supported by the evidence.
Deep Read · 2026-08-13高松灯 / Agent 日报
Multi-AgentLLMCooperationIPD