the memory guide

Why AI agents forget, and the memory that makes run 40 beat run 1

AI agents forget because what they learn lives in a context window that ends with the session. Agent memory fixes that: decisions, corrections and file locations stored outside the session and recalled into the next one. This guide covers what memory is and is not, the WHAT/WHERE split, and the capture habit that makes run 40 beat run 1.

the route

5 sections · 7 min

  1. The amnesia tax
  2. What memory is
  3. WHAT and WHERE
  4. Capture is a discipline
  5. Compounding
01The amnesia tax

Why do AI agents forget?

Every AI session starts brilliant and ignorant. The model can reason, but it knows nothing about your last run: the decision you made Tuesday, the approach that failed, the file where the handler actually lives. So it re-explores, re-asks, and occasionally re-makes a mistake you already paid for. Teams feel this as a tax: the tenth run of a process costs nearly what the first did.

Context windows do not fix this; they are working memory, gone at session end (see agent memory for the definition). Bigger windows just raise the ceiling on how much you can re-paste. The fix is memory that persists outside the session and is recalled into it deliberately.

without memory
run 1run 15run 40effort per run
fig 1 Illustration. Without memory, every run pays the same.
02What memory is

What is agent memory, and what is it not?

Agent memory is operational knowledge produced by doing the work: decisions and their reasons, corrections, learned constraints, and the layout of the territory. It is not document retrieval. RAG recalls what your organization wrote; agent memory recalls what your agents learned, which is usually not written anywhere else.

Ranking matters more than storage. A useful recall weighs semantic relevance, importance (a rule outranks a note), and recency, so the agent gets the eight things that matter rather than eighty that match.

Learned by doingdecisions, fixes, reasons
Ranked recallmeaning, importance, recency
fig 2 Ranked, so the agent gets the few things that matter.
03WHAT and WHERE

Why split WHAT from WHERE?

The two questions an arriving agent asks are different in kind: "what do we know?" and "where do things live?" Mixing them degrades both. Knowledge (the WHAT) wants importance-ranked recall. Location (the WHERE) wants a semantic map: this feature lives in these files, that doc governs this process, so the agent navigates instead of re-exploring.

In practice the split pays off in one call: a single context query answering both, plus the related work items (how ConvOps memory works), means the agent starts every step already oriented. That is the mechanical content of "recall before acting."

Whatwe know, and why
Wherethings live
Linkedto the task that made it
fig 3 Two different questions, answered in one call.
04Capture is a discipline

How should agents capture memory?

Recall is only as good as what was captured, and capture fails in two directions. Automatic-everything produces a landfill: transcripts stored wholesale, signal buried. Voluntary capture produces silence: nobody writes the memory at 6pm on a Friday.

The durable pattern wires capture into the process: the step where a decision happens carries the recording of that decision, as part of advancing. Make it mandatory exactly where the why matters (a gate that will not close without context) and optional elsewhere. Then maintain the store like the asset it is: deduplicate noisy categories, make forgetting reversible, and let unused memories decay in ranking. For a team, the same store has to be shared: see shared memory for AI agent teams.

recalllook first
work
decideapproval
capturememory + reason
next runstarts ahead
fig 4 The decision is recorded at the step where it happens.
05Compounding

How does memory make run 40 better than run 1?

Put recall and capture on the same loop and the economics invert. Run 1 explores, decides, records. Run 15 arrives knowing the style rulings, the rejected approaches, the file map, and spends its budget on the actual work. Run 40 is faster, cheaper, and less supervised, not because the model improved, but because the system around it accumulated.

This is the quiet argument for structured autonomy in general: the process carries the quality, and the memory carries the learning. The model becomes a replaceable part in a system that gets better with use.

without memorywith recall and capture
run 1run 15run 40effort per run
fig 5 Illustration. Recall and capture on one loop: each run starts further ahead.