Skip to the document
Madhuopen lab
Why AI Agents ForgetPart I — Why memory, and what kinds

Chapter 2

What kinds of memory does an agent need?

The two big kinds of memory an AI agent has, short-term and long-term, and the three flavours of long-term memory. We will also see why the same fact belongs in a different place depending on how it will be used.

4 min read825 words3 recall cards

Before you read, guess

How does the text describe long-term memory?

Take ten seconds and guess — even a wrong guess makes the answer stick. Tap to see where the chapter lands, or just read on.

Long-term memory is the card file, and every card is a fact, an event, or a rule.

In this blog, we will learn the two big kinds of memory an AI agent has, short-term and long-term, and the three flavours of long-term memory. We will also see why the same fact belongs in a different place depending on how it will be used.

Short-term memory is the context window

Short-term memory is what the model can see right now: the live turns of this conversation. It is the sticky note on the monitor. Cheap to write, always in view, and gone when the session ends.

turn  1   [system | memory | t1                                   ]   ~250 tok
turn  6   [system | memory | t1 t2 t3 t4 t5 t6                    ] ~1,500 tok
turn 40   [system | memory | ....... 40 x 250 = 10,000 tok ...... ]
                            |
                            v  summarize turns 1 to 35, keep the last 5
turn 40'  [system | memory | summary ~150 tok | t36 t37 t38 t39 t40]

When a single session gets long, we compress the old turns into a summary and keep the newest turns word for word. This is a rolling summary.

One catch: summaries are lossy. The class experiment shows a 14-message trip-planning chat compressed 6.5x into 84 tokens, and the exact "120 euros a night" happened to survive. It is not guaranteed to. Any detail can vanish in a compression pass. That is why long-term memory stores discrete facts, not summaries.

Long-term memory survives the session

Long-term memory is the card file from the last blog. It lives in a database, not in the prompt. It has three flavours, and the names come straight from how psychologists describe human memory.

            the memory of an agent
            +----------------------------------------------+
 SHORT-TERM | working memory = the context window          |
 (this      |   the live turns of this conversation        |
  session)  |   + a rolling summary when it overflows      |
            +----------------------------------------------+
 LONG-TERM  | semantic    facts:   "Maya is vegetarian"    |
 (across    | episodic    events:  "last week we planned   |
  sessions) |                       the Lisbon trip"       |
            | procedural  how-to:  "answer short and       |
            |                       direct"                |
            +----------------------------------------------+

Semantic memory: what is true

Semantic memories are facts about the world and the user. "Maya is vegetarian." "Maya's partner Sam has a birthday on June 18." They tend to stay true for a long time, and when they change, the old one should be replaced, not stacked next to the new one.

Analogy: the "About me" page. You do not add a second address when you move. You change the address, and maybe keep a note of the old one.

Episodic memory: what happened

Episodic memories are events with a time. "Maya planned a Lisbon trip the week of June 21, 2026." "Maya's dog Poppy had a vet checkup on March 14." They are useful for a while and then decay into noise. A dentist appointment from two years ago is not worth a slot in the prompt.

Analogy: a diary. Every entry is dated. You rarely reread the old ones, but you never rewrite them.

Procedural memory: how to behave

Procedural memories are standing instructions. "Answer short and direct." "Use metric units." They are not facts about the world; they are rules for the assistant. They are small, they never expire on their own, and the user should be able to edit them like settings.

Analogy: a habit. You do not think about it each time; it shapes every answer.

Why the split matters

The flavour decides the policy. This one table is most of the design:

CategoryExpire?On changeInjected when
semanticneversupersede the old factrelevant to the question
episodic~90 days after last useappend a new eventrelevant, with its date
proceduralnever, user-editablereplacealways, they are cheap

Get the category wrong and you get the wrong policy. A birthday is a recurring date, so it is semantic, not episodic. A one-time trip is episodic. "Keep answers short" is procedural even though it sounds like a preference.

Where a rolling profile fits

Some systems keep a fourth thing: a short running profile of the user, a paragraph that says who they are. The BEAM paper's LIGHT framework calls this a scratchpad. It is not a replacement for the card file. It is a cheat sheet for the extractor, so that "she" in a new message resolves to "Maya" without a lookup. We treat it as optional and measure whether it pays for itself.

Closing

Short-term memory is the sticky note you throw away tonight. Long-term memory is the card file, and every card is a fact, an event, or a rule. The label on the card decides how long it lives and when it gets read.

Before you go

In one sentence, what was this chapter about?

From memory, without scrolling up. Writing it is what makes it yours; the grade is only to show you what you had.

How sure?