SeoWeb
  • Work
  • Services
  • CV
  • Contact
    • AI
  1. Home/
  2. AI Handbook/
  3. Agents & Evaluation/
  4. Long-term memory and state management

[ Agents & Evaluation ]

Long-term memory and state management

Target audience: technical + deep-diving non-technical | Prerequisites: 4.2 RAG, 3.2 Context management

What you'll learn

After this document you will be able to:

  • distinguish conversation-scoped memory from long-term memory (information across sessions) and explain why both live in the system's data, not in the model;
  • decide which facts are worth keeping in memory and which are not;
  • build a simple customer profile (a collection of structured persistent facts) and include its important fields with every request;
  • plan memory updating, forgetting, and human oversight of changes;
  • keep state (where a process stands) consistently, and know the two main risks: stale information and context poisoning.

In plain terms

The model never remembers you — your system does. When a customer comes back days later, the model knows only what the system writes for it from its notepad for that call. Long-term memory is that notepad: only important persistent facts go into it; they are overwritten when the customer corrects them, and they are included at the start of every new call.

Two kinds of memory: the conversation inside and between sessions

From document 3.2 you know the basic truth: the model has no memory; every request is the whole conversation from the beginning, and “remembering” is the system's job. There one conversation was examined; here we look across its boundary.

Conversation-scoped memory. During one session (a connection period), the system keeps the conversation history and the “state of affairs” (a record of the important facts, the 3.2 principle). This memory's lifetime ends with the session: the call ended, the window emptied, and something remained only if the system itself wrote it down.

Long-term memory is information across sessions: the customer's preferences, past cases, agreements. A customer comes back days later and the system must know them — not because the model remembers anything, but because the system keeps the facts in its own database and puts them into the context window (the amount of text the model sees at once) before every new session.

Neither kind of memory thus lives in the model — both are system data carried into the window. The difference is only in lifetime: conversation-scoped memory lasts one session, long-term memory years.

What to keep in memory and what not

In plain terms: memory is not storage where everything fits — it is a selection of what will help in the next call. The fewer entries, the more certain that the important doesn't get lost under the noise.

Three kinds of entries are worth keeping:

  1. Persistent facts that rarely change: the customer's number, language preference, the managed device.
  2. Agreements that both parties have clearly discussed: “wants email notifications”, “doesn't want advertising”, “may not answer the phone in the evenings”.
  3. Important historical cases that affect future service: “on September 28 the channel was switched, the problem remained” — the next call must know this.

Not worth keeping:

  • Everyday noise. The verbatim text of every call is not memory but a log. If everything is written down, memory grows into noise that pushes the important down.
  • Sensitive data without a reason. For every fact ask: does it help give a better answer in the next call? If not, don't store it; 3.7 Security teaches how to reduce and protect data.

How to store: profile before complexity

In plain terms: a customer profile is like a doctor's appointment card — at the start of every visit the card is reviewed, rather than asking for the patient's life story from scratch. The card is short, structured, and carefully maintained; exactly that kind is enough.

Start with the simplest: a structured customer profile in a database or even a file — with fixed fields (language, preferences, notes), not a pile of free text. The system puts the profile's important fields into the context of every request: this is the long-term version of the “state of affairs”, continuing where 3.2 left off. The example with Mart below shows what such a request looks like.

This is enough in most cases. A bigger solution enters the picture when there are thousands of past cases and they don't fit into a single profile: then the system keeps the cases separately and retrieves the relevant ones based on the new question — this is retrieval-augmented generation, RAG (see 4.2).

State: where the process stands

Alongside facts, the system keeps state. An order has a state: new → shipped → returned. If state gets tangled, the customer is asked again or an action is done twice.

Across sessions, keeping state is especially important for flows left half-finished: the customer started a return on Monday, filled the form halfway, and came back on Thursday. If the system keeps the flow's state, it continues where it left off; if it doesn't, the customer must do everything again. The rule is simple: every multi-step flow has its own state, which at any moment has one valid value — and that value is kept consistently in the system's data, not in “the model's memory”, which doesn't exist.

Updating, forgetting, and the risks

In plain terms: memory is not truth but a hypothesis. A recorded fact may turn out wrong tomorrow — the person changes their mind, the device is replaced, the agreement ends. So: update the entries, confirm the changes, and double-check when it matters.

Updating. Entries age. When the customer reports a new preference, the system overwrites the old entry — it doesn't add a new one alongside. Two conflicting entries are worse than one outdated one: the system no longer knows which holds.

Who may change. A memory change must not happen silently: the more important the entry, the more clearly a human must confirm — or at least see — the change. Human-in-the-loop (a workflow where a human approves the result before it is used) makes sense for memory precisely at changes (see 3.5 Safety). The usual split: lighter facts (the device model) update automatically, more sensitive ones (contact details, agreements) go to a human for approval.

Forgetting. A customer may say “forget me” — and the system must be able to delete their entries. The basis for this is the customer's right to their data and data protection requirements (see 5.3 Security and data protection).

Two risks to keep in mind:

  1. Old information goes wrong. The “truth” recorded in memory can be old: the person changed their mind. Long-term facts are therefore not commands to execute for the answer, but background to double-check when it matters: “Do you still want notifications by email?”
  2. Context poisoning. If a couple of wrong entries get into the window, they poison all subsequent answers — the model confidently relies on the wrong fact at every request. Faulty memory is thus even more dangerous than a hallucination (the model's confident but unfounded answer, see 1.1): the source of the error is not the model here but your own data, and the error repeats systematically.

A step-by-step example: Mart's third call

Situation: the CallHelp phone support (ordering the example with 2.4). The customer Mart calls for the third time about the same problem: the internet drops out every couple of days.

a) Without long-term memory (before). The system has nothing about Mart; every call starts from zero:

“Hello! Please give your customer number and describe the problem.”

Mart explains the same story a third time: which modem, when the internet drops, what has already been tried. To the customer it feels like talking to a wall — a bad experience that is useful to no one.

b) With a customer profile (after). The system identifies the profile by the number and puts the important fields at the start of the request in the form of a “state of affairs”:

{
  "state_of_affairs": {
    "customer": "Mart Tamm (Customer #207)",
    "language": "et",
    "modem": { "model": "KTX-200", "year": 2023 },
    "previous_calls": [
      { "date": "2026-09-18", "topic": "internet drops in the evenings", "resolution": "modem restart" },
      { "date": "2026-09-28", "topic": "internet drops again", "resolution": "channel switch, problem remained" }
    ],
    "preferred_contact": "email",
    "notes": "third call on the same topic — raise priority"
  },
  "new_message": "This internet again! Third time now."
}

The answer:

“Welcome back, Mart! I see this is the third call about the internet dropping — on September 18 a modem restart helped, on September 28 the problem remained despite a channel switch. Your KTX-200 is a 2023 device; I now suspect the modem's age. I'll send an offer by email to swap the device for a new one — as agreed, we'll stay with email as the contact method.”

All the facts came from the profile — not from Mart's memory or the model's luck.

c) Updating. Mart says he now has a new modem, an NXR-50. The system overwrites the entry:

{ "modem": { "model": "NXR-50", "year": 2026, "updated": "2026-10-05" } }

Note: no trace of the old model remains in the profile. If the KTX-200 entry were left in place, the system would assume in the next call too that the old device is the subject — exactly the context poisoning described above.

Summary

  • Two kinds of memory: conversation-scoped memory (history and the “state of affairs” during one session) and long-term memory — information across sessions that the system keeps in its own data and puts into the window at the start of every request.
  • Keep only what matters: persistent facts, agreements, and important historical cases; everyday noise and sensitive data without a reason stay out.
  • The simple solution first: a structured customer profile with fixed fields; if there are many cases, RAG comes to help (4.2).
  • State is kept consistently: a multi-step flow continues where it left off — even across sessions.
  • Memory needs management: entries are overwritten, not accumulated; the more important changes are approved by a human; the customer can demand to be forgotten.
  • Memory is a hypothesis, not truth: an old entry may be wrong, and a couple of wrong entries poison all subsequent answers — so update memory, double-check, and verify.

What's next?

  • previous → 4.2 RAG
  • next → 4.4 Multi-agent architectures
  • Managing conversation-scoped context → 3.2 Context management
  • back → handbook index

Last updated 2026-10-05

← PreviousRAG: using your own data as a source of answersNext →Multi-agent architectures

© 2026 Siim Liimand · SeoWeb

GitHub/AI Handbook/Tallinn, Estonia

59.4370° N, 24.7536° E — Tallinn, Estonia

↑ Top