章节结构(全文标题提取): 1 Introduction 2 Related Work / The Retrieval Paradigm and Its Limits 3 The Compiled Memory Position 4 A Reference Implementation 5 Empirical Evidence 6 Discussion and Limitations 7 Conclusion and Future Work / Implications and Open Problems Acknowledgments Declaration on Generative AI · 2.1 Conversational Memory, Reasoning Memory, and Personalization · 2.2 Skill Libraries and Agentic Skills · 2.3 Multi-Agent Systems · 2.4 LLM-as-Judge Evaluation · 3.1 Three principles · 3.2 Compiled vs. retrieved: where they diverge · 3.3 When compilation is the right choice · 4.1 Phase 1: History Harvest · 4.2 Phase 2: Pattern Analysis and Swarm Generation · 4.3 Phase 3: Runtime Augmentation
Memory for LLM agents has converged on a single architectural pattern: store experience as text, embeddings, reflections, or rules; retrieve at inference time; let a general-purpose orchestrator interpret what to do. This paper argues that the pattern is the wrong default for personalization. We position Muscle Memory - the practice of compiling recurring user intent into purpose-built specialist agents - as a distinct memory paradigm from retrieval, and we argue that compilation is a better fit for the workloads where current assistants impose a multi-turn tax on their users: making them repeatedly correct format, depth, and scope to obtain a domain-appropriate answer. We support the position with a reference implementation and empirical evidence. The implementation is a four-phase pipeline (Harvest $\rightarrow$ Analyze $\rightarrow$ Augment $\rightarrow$ Evaluate) that mines conversational history, separates behavioral from task patterns, and emits quality-gated executable compiled specialists with two-stage trigger matching. On 90 held-out scenarios across five user personas, the augmented assistant wins 32 of 36 cases where a specialist fires, an 88.9% win rate, with a +2.05 personalization gain and only a $-0.28$ accuracy cost on a 1-4 scale. We discuss why compilation is better suited than retrieval in this regime, what the result implies for the broader memory design space, and what open problems remain.