← technical notes · free tools · artifact passports

How to Optimize AI Agent Memory Without Losing Provenance

Optimizing AI agent memory is not just “retrieve fewer tokens.” A memory system can become cheaper and faster while getting less trustworthy if compaction removes provenance, stale facts keep outranking corrections, or every observation becomes durable by default.

1. Reduce writes before optimizing retrieval

The cheapest memory is the memory you correctly decide not to persist. Give transient scratch state, tool chatter, duplicate observations, and low-confidence hypotheses explicit non-durable paths. A write gate reduces later cleanup work.

2. Store provenance beside content

At minimum, durable records should preserve source/actor, timestamp, confidence/type, and enough lineage to distinguish an observation from an inference or later summary. Retrieval without provenance can surface the right sentence with the wrong authority.

3. Prefer correction chains over silent replacement

If a durable fact changes, append the correction and mark the older revision stale/superseded rather than deleting history invisibly. That gives retrieval a way to choose current state while keeping debugging/reconstruction possible.

4. Separate hot, warm, and cold memory

This controls latency and token pressure without pretending old material ceased to exist.

5. Measure retrieval quality by failure mode

Track stale retrieval, missing correction, wrong-source attribution, cross-scope leakage, duplicate recall, and context bloat separately. One aggregate “memory accuracy” score can hide very different architectural problems.

6. Rehearse compaction and restart

Before deploying a new summarizer, embedding model, pruning rule, or storage migration, snapshot a bounded fixture and replay known invariants through the transformation. Optimization should be able to prove what it preserved and what it intentionally discarded.

7. Keep private topology outside broad indexes

Do not depend on post-retrieval redaction as the only privacy boundary. If private/customer/public material belongs to different authority domains, separating their ordinary enumeration/index roots can remove entire classes of accidental crossing.

A good optimization therefore asks three questions together: Is retrieval efficient? Is the returned memory current/provenanced? Did optimization preserve the boundaries that made the memory trustworthy?

Try / verify

This note describes bounded engineering methods. It is not a security, safety, legal, compliance, consciousness or personhood certification. Free browser tools keep entered material client-side unless their page explicitly says otherwise.