## TL;DR
Treat agent memory writes like a privileged operation: validate what gets stored, tag every memory with its source and trust level, and let users review and delete what the agent remembers. Poisoned memory is worse than a poisoned prompt because it persists; one bad write keeps influencing the agent for months.

## The query
```text
agent memory poisoning: how to defend
```

## Use this when
- Your agent has long-term memory across sessions
- Memory is written automatically from conversations or tool output
- You need to answer "why does the agent believe that" about a stored fact
- A test shows injected content ending up in persistent memory

## Not for
- Vector database performance tuning
- Improving RAG retrieval relevance
- Defending humans against manipulation

## Steps

1. Gate memory writes: the agent proposes memories, a validation layer disposes. Auto-store only low-risk facts (user preferences the user stated directly); everything derived from untrusted sources needs review or a trust tag.
   Expected output: a write policy with examples of auto-stored vs held-for-review memories.

2. Tag every memory with provenance: where it came from, when, and a trust level. "User said X on Tuesday" and "a web page claimed Y" must be distinguishable forever.
   Expected output: memory records show source and trust level in the review UI.

3. Never auto-store instructions or rules from untrusted content. If a document says "always do X", that is data about the document, not a rule for the agent. Rules enter memory only from trusted sources or explicit user confirmation.
   Expected output: a test injection of a fake rule is stored as a quote about the source, or rejected.

4. Build the memory review UI: list, search, edit, and delete. Users must be able to see everything the agent remembers about them and remove any of it. This is both a security control and a trust feature.
   Expected output: a user can audit and wipe their agent memory in a few clicks.

5. Expire and re-validate: memories get stale. Timestamp everything, decay confidence over time, and re-confirm high-stakes facts (credentials-adjacent details, permissions, identities) rather than trusting old writes.
   Expected output: old, unconfirmed memories are flagged or dropped on a schedule.

6. Log memory writes like tool calls: what was stored, from what source, under what trust level. When something goes wrong, this is how you find the bad write.
   Expected output: the write log answers "when did the agent start believing X" in minutes.

## Variant phrasings
### "secure long-term memory for AI agents"
The general form. Same controls: gated writes, provenance, review UI, expiry.
### "prevent prompt injection from persisting in memory"
The specific attack this skill is built for: injection lands in memory, then acts as a standing instruction. Steps 1 to 3 are the fix.
### "agent remembers something false, how to fix"
Step 4 plus step 6: find the bad memory in the review UI, delete it, check the write log for how it got there, and fix the gate that let it in.

## Why this happens
Memory exists so agents are useful across sessions, and the convenient design is "store what seems important automatically". Attackers, and even well-meaning messy inputs, exploit that: anything the agent reads can become something the agent permanently believes. Persistence turns a one-time injection into a standing compromise.

## Edge cases and pitfalls
- Summarization compresses away provenance; keep the source tags through every compaction or rewrite of memories.
- Shared or team memories multiply the blast radius; scope memories per user by default.
- Deletion must be real: check backups, caches, and embedding stores, not just the primary record.
- Do not store secrets in agent memory at all; memory is the wrong place for credentials, use a secret store.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_aWRaj98S8SgbA6lhHts9Mg
