The Hook: Upgrading Should Not Mean Forgetting
Your AI agent has been running for six months. It remembers your preferences, your project history, your team's decisions. Then you upgrade the underlying model from version 4 to version 5. The agent still runs. It still responds. But it has quietly forgotten half of what it knew.
No error message. No crash. Just wrong answers where there used to be right ones.
This is the problem that Ankit Goyal and Jaideep Ray at LinkedIn set out to measure in their paper "Does Your Agent's Memory Survive a Model Upgrade?" published September 4, 2026. They tested four different memory storage formats under controlled model swaps, and the results reveal a landscape of silent failures that most teams are not testing for.
Why This Matters Now
Every major agent framework today offers some form of persistent memory. LangChain has memory modules. AutoGen stores conversation histories. Custom agents compress past interactions into summaries or knowledge graphs. The assumption is that this memory persists across upgrades.
But memory is not just storage. It is a contract between the writer (the model that created the memory) and the reader (the model that consumes it). When you swap the reader, that contract can break in ways that produce no errors but degrade accuracy significantly. The paper quantifies exactly how much.
The Four Memory Formats
The study compares four distinct approaches to agent memory, each representing a real-world pattern used in production systems:
Each format represents a different trade-off between storage cost, retrieval speed, and information density. What the paper adds is a fourth dimension: portability, the ability to survive a model change without accuracy loss.
The Experiment: A Rigorous Setup
The experimental design is notably careful, following pre-registered hypotheses with a signed Git tag to prevent p-hacking. Here is what they built:
The two models are Llama-3.1-8B-Instruct and Qwen2.5-7B-Instruct-1M, both similarly sized but architecturally distinct. Questions cover direct facts, temporal changes, contradictions, multi-step relations, and aliases. Crucially, answers use randomized codes to eliminate guessing and make exact-match scoring reliable.
Finding 1: Fixed Schemas Transfer, Compressed Notes Do Not
The headline result is stark. When you swap the model that writes memory with a different model that reads it, fixed-schema knowledge graphs barely notice. Compressed notes can swing by over 13 percentage points.
| Format | Reader | Own-Store Accuracy | Inherited Accuracy | Delta |
|---|---|---|---|---|
| KG-fixed | Llama | 0.8456 | 0.8445 | -0.11 pp |
| KG-fixed | Qwen | 0.9878 | 0.9880 | +0.02 pp |
| NOTES | Llama | 0.3762 | 0.4753 | +9.91 pp |
| NOTES | Qwen | 0.4719 | 0.3391 | -13.28 pp |
KG-fixed shows a combined shift of +0.0004 with standard error 0.0020 across both migration directions. That is statistical noise. The schema enforces a contract that both models can read regardless of who wrote it.
NOTES, by contrast, show a 23-percentage-point spread depending on direction. When Llama reads Qwen's notes, accuracy jumps by 9.91 points. When Qwen reads Llama's notes, accuracy drops by 13.28 points. Same format, opposite outcomes. Averaging these two numbers gives a misleading -1.69 points that hides the real variance.
Finding 2: Mixed Embeddings Are Worse Than Either Extreme
For RAG systems, model upgrades often mean new embedding models. The natural migration strategy is incremental: keep old embeddings, add new ones as new data arrives. The paper tested this with BAAI/bge-large-en v1.0 and v1.5 (both 1024-dimensional).
Full re-embedding gains 11.90 percentage points over the old index. The 50/50 mixed index recovers only 4.96 points of that gain, forfeiting nearly 60% of the improvement. No errors. No warnings. The system just retrieves worse chunks because embeddings from different model versions occupy slightly different regions of the vector space.
The ideal routing upper bound (59.60%, a hypothetical where each query is routed to whichever embedding version works better) shows that the information is there. The mixed index just cannot find it.
Finding 3: Same Symptom, Different Disease
One of the paper's most useful contributions is a diagnostic decomposition that isolates where information is lost. NOTES and RAG both lose accuracy compared to raw transcripts, but for entirely different reasons.
For NOTES, 80% of accuracy loss happens during construction, when the model compresses the conversation into a summary. Information is discarded at write time and cannot be recovered at read time. The summary itself is the bottleneck.
For RAG, 81% of accuracy loss happens during retrieval. The information exists in the chunks, but the wrong chunks are returned. The retrieval pipeline, not the storage format, is the bottleneck. Construction loss is only 1%.
Finding 4: Raw History Is Your Insurance Policy
Can you repair broken memory after a migration? The paper tests two repair strategies: store-only (work with what you have) and raw-retained (re-derive from original transcripts).
Raw-Retained Repair
- NOTES (Qwen repairing Llama): 34/48 cases hit 90% target
- RAG re-embedding: 48/48 at 90-99% targets
- KG-fixed rebuild: 48/48 at 90%, 45-46/48 at 99%
- Median cost for NOTES repair: $0.76
- Median cost for RAG re-embed: $0.013
Store-Only Repair
- NOTES repair success at 90% target: 0 out of 48
- No raw history means no re-derivation possible
- Compressed summaries cannot be uncompressed
- Information lost at construction is gone permanently
The numbers are unambiguous. Without raw transcripts, NOTES repair succeeds exactly zero times out of 48 attempts. With raw transcripts, 34 out of 48 cases reach 90% recovery at a median cost of $0.76. The raw history is not just archival, it is operational infrastructure.
Finding 5: Direction Matters
Perhaps the most counterintuitive result is that migration is not symmetric. Swapping Model A for Model B produces different outcomes than swapping Model B for Model A.
Llama reading Qwen's NOTES gains 9.91 percentage points. Qwen reading Llama's NOTES loses 13.28 points. These are not minor variations. They are directionally opposite effects that would cancel to a misleading average of -1.69 if you did not measure them separately.
This has direct implications for upgrade testing. If you validate a migration from GPT-4 to GPT-5, that validation does not tell you what happens when you roll back from GPT-5 to GPT-4. Each direction must be tested independently.
What This Means for Practitioners
The paper distills into five operational rules for anyone building agent memory systems:
- Keep raw conversation transcripts. Compressed memories are disposable derivatives, not the source of truth. Raw history enables recovery. Without it, you have no rollback path.
- Prefer fixed-schema memory formats. Knowledge graphs with rigid schemas transfer across models with near-zero accuracy loss. The schema is the contract. Free-form notes are the most fragile format tested.
- Test each migration direction separately. Averaging bidirectional results conceals real failures. Validate A-to-B and B-to-A independently.
- Commit to full re-embedding or accept the cost. Mixed embedding indices (old + new) recover less than half the potential improvement. There is no cheap middle ground that preserves accuracy.
- Diagnose before fixing. NOTES failures are construction problems (fix by re-deriving). RAG failures are retrieval problems (fix by re-embedding). Same symptom, different treatment.
My Take
This paper fills a gap that practitioners know exists but rarely measure. Everyone who has upgraded a model in a production agent system has encountered unexplained behavior changes. This study puts numbers on the problem and reveals a clear hierarchy: fixed schemas are safe, RAG is fixable, and NOTES are fragile.
The strongest contribution is the diagnostic decomposition. Knowing that NOTES lose information at construction (80%) while RAG loses it at retrieval (81%) transforms the debugging process from guesswork into targeted intervention. This is the kind of result that changes how you build systems, not just how you think about them.
The limitation is scope: two similarly-sized open-weight models, synthetic histories, and single-stage dense retrieval. Real production systems involve larger models, multi-stage retrieval, and subjective conversations. The directional asymmetry finding suggests that scaling up will reveal more surprises, not fewer.
For anyone building agent systems today, the practical message is clear: treat memory as infrastructure, not an afterthought. Model upgrades are inevitable. Memory migration plans should be part of the architecture from day one, not retrofitted after the first silent failure.
Discussion Questions
Open questions for the weekly discussion. No clean answers here.
Memory is not a feature. It is infrastructure. The models will keep changing. The question is whether your memory system was built to survive that.
Read the Full Paper →