SeoWeb
  • Work
  • Services
  • CV
  • Contact
    • AI
  1. Home/
  2. AI Handbook/
  3. System Architecture/
  4. Context management: how the model “remembers”

[ System Architecture ]

Context management: how the model “remembers”

Target audience: technical + non-technical readers who want to go deeper | Prerequisites: 3.1 API integrations

What you'll learn

After this document you will be able to:

  • explain why the AI model has no memory and what “remembering” means in that case;
  • see how conversation history grows and why a long history makes the system expensive and forgetful;
  • choose and combine three strategies: a limited window, a summary, a “state of affairs”;
  • place the system prompt's instructions in the right spot in the context;
  • say when the knowledge doesn't fit in the window and when RAG and long-term memory come to the rescue.

In plain terms

An AI model remembers nothing. Every request is a new conversation from zero for it — about the past it knows only what the system writes out for it anew each time. “Remembering” is therefore not the model's but the system's work: before every request, put into the window everything a good answer needs — instructions, key facts, earlier turns — and nothing else.

The model has no memory — all information must be sent along

From 1.1 you know what a context window is (the amount of text the model sees at one time). Here you need one of its consequences:

The model has no memory of the conversation. Every API request is the whole conversation from the beginning.

If a customer wrote something yesterday and today sends a new message, the model knows nothing about the first — unless the system sends yesterday's text along in the request. Even chat applications' “memory” is exactly this: the system keeps the earlier messages in its own data and puts them back into the window every time.

Conversation history and why it must not grow unchecked

In plain terms: conversation history is like the minutes of a meeting that are read out in full at every new agenda item — the longer the meeting, the longer the reading every time, and the greater the chance that the middle pages go unnoticed.

Conversation history (the collection of earlier turns that the system sends along again with every request) is the simplest way to give the model “memory”. A turn (one round of the conversation — the user's question + the model's answer) is appended to the end of the history, and on every subsequent request the whole history goes along again.

Three reasons why this must not be allowed to grow unchecked:

  1. Costs grow with every turn. Every token is billable and a long history is spent again on every request. How to account for and limit costs, 3.6 Cost management teaches.
  2. Attention is uneven. The middle part of a long input gets the least attention: an order number, said at the start of the conversation and now in the middle of a long history, is exactly the information most easily lost (see 1.3 Prompt fundamentals).
  3. The window fills up. When the history exceeds the window's capacity, something must be left out — either the system cuts off the beginning mechanically, or you decide yourself what survives. The second option is always better.

Three strategies for managing a long conversation

In plain terms: the window filling up resembles the papers piling up on a table. Three actions: (a) throw the older ones away and keep the latest, (b) put the older content on one summary sheet, (c) keep the important things in a separate notebook.

1. A limited window

Keep only the last N turns of the history (for example ten) — the system cuts the older ones out. The simplest strategy: one limit number in the settings, no extra work.

The risk is direct: everything said early on disappears by default. An order number, said in the first turn, is gone after ten turns — the model either asks again or, in the worst case, figures it out itself (a hallucination — the model's confidently stated but wrong answer).

2. A summary

Summarization (a short summary of the older part of the conversation that the system has the model produce) keeps the history from growing constantly: when there are too many turns, the system replaces the older part with one short summary. From then on, the summary + the recent turns go into the window.

Early information survives in the summary, but a summary is always a selection — the model may leave out exactly the detail that later turns out to be important (for example the order number).

3. A “state of affairs” — separating out the key facts

The system keeps the important facts as a separate record alongside the history — the “state of affairs” (a record of key facts): who the customer is, which order, what the goal, what has been decided. The system doesn't let the model summarize these facts or leave them to the history's care — they come from the system's own data and go along with every request. The shape of one request:

{
  "system_prompt": "You are the customer service assistant of an e-store. Answer briefly and in English.",
  "state_of_affairs": {
    "customer": "Kadri Kask",
    "order": "1224",
    "product": "table lamp Nordica",
    "wish": "exchange the black one for the dark green one"
  },
  "recent_turns": [
    { "role": "user", "text": "Is the exchange still possible?" },
    { "role": "model", "text": "Yes, within 14 days of purchase the exchange is possible." }
  ],
  "new_message": "Like I said, I would still like a different color."
}
StrategyHowAdvantageRisk
Limited windowKeep only the last N turnsSimplest, no extra workEarly information disappears by default
SummarySummarize the older part; keep summary + recent turnsHistory survives, the window doesn't grow constantlyThe summary may leave out an important detail
“State of affairs”Key facts as a separate record, sent with every requestFacts accurate and always presentNeeds upkeep: who and when updates the facts

In practice they are combined: the system keeps the “state of affairs” + the recent turns and occasionally has the older part summarized.

When the knowledge is bigger than the window

The three strategies assume the needed information exists somewhere — the only question is what goes into the window. When the knowledge itself is bigger than the window (e.g. a whole company's document base), RAG helps: the system searches out the relevant passages based on the question and gives them to the model (see 4.2 RAG). When information must be remembered across sessions — the customer comes back a week later and the system should recognize them — the facts are saved into the system's data and brought back into context as needed; this is called long-term memory (see 4.3 Long-term memory and state management).

The system prompt: the rules that always stand in front

In plain terms: the system prompt is like the job manual given on the first day of work — the rules written down once that the assistant keeps in mind on every task.

The system prompt (the standing instruction that is always in front, defining the role and the rules) is the context's first and most permanent part: it doesn't change with every turn but holds for the whole conversation. It contains the role (“You are the customer service assistant of an e-store”), the rules (“answer briefly, don't invent prices”), and the limits (what the assistant may not decide).

Its place is no accident — a typical request is built up like this:

  1. the system prompt — role and rules, because they always hold;
  2. the “state of affairs” and the relevant data;
  3. the history or the recent turns;
  4. the new question — last, because that's what gets answered.

This keeps the important information at the beginning and end of the window, where attention is strongest (see 2.4 Inputs and data preparation).

One warning: the standing instruction is part of the prompt (the instruction given to the model) and spends tokens on every request. So keep it precise and short — every sentence put there is paid for anew each time.

Step-by-step example: the customer comes back the next day

The situation: an e-store's customer chat. Yesterday the customer went through a conversation around order 1224 (the table lamp Nordica) — over 12 turns it came out that the customer wants dark green instead of black. Overnight the conversation ended. The next day the customer opens the chat window again:

“Like I said, I would still like a different color.”

a) Without context management (before). The new conversation consists only of this sentence. The model doesn't know which product is meant, what the order is, or what color the customer has in mind:

“We'll be happy to help! Please tell us which product and which color you mean.”

The result is bad: the customer has to explain everything again — or, in the worst case, gets a plausible-sounding but irrelevant answer.

b) A limited window of 10 turns. The system keeps the last 10 turns, but yesterday's conversation was 12: the first two — exactly the ones where the order and the desired color were discussed — fell outside the window. The model does see that the conversation concerns an exchange, but no longer knows the product or the color:

“Of course! Please tell me which product you mean and what color you would like — I'll check right away whether an exchange is possible.”

The information exists in the system's data, but it didn't fit into the window — early information was left out.

c) With a “state of affairs” solution (after). The system keeps the record of facts compiled during yesterday's conversation (product, order, wish) and sends it along with every request — in the JSON shape shown above. The new question reaches the model together with all the facts:

“Welcome back! Let's continue where we left off: the order 1224 table lamp Nordica can be exchanged from black to dark green, and the exchange is free. Please confirm, and the new lamp will go out right away.”

Before, the customer had to repeat all of yesterday's information; after, the answer is personal and correct — the facts came from the system's record, not from the model's luck.

Note: option c didn't need the long history at all — the whole “memory” is a short record of facts that the system keeps. The same holds for workflows: every workflow run is independent, and its “memory” comes from the data (see 2.3 Basic workflows).

Summary

  • The model has no memory — every request is the whole conversation from the beginning. “Remembering” is the system's work: putting into the window everything a good answer needs.
  • Conversation history grows with every turn and goes along in full every time — that's why costs and the risk of lost information grow too.
  • Three strategies: a limited window (simple, but early information is lost), a summary (history survives in shortened form) and a “state of affairs” (facts accurate and always present). In practice they are combined.
  • The system prompt is the context's most permanent part: role and rules first, then facts, then history, and the new question last.
  • When the knowledge is bigger than the window, search and memory help: RAG for large document bases, long-term memory across sessions.

What's next?

  • previous → 3.1 API integrations
  • next → 3.3 Tools and actions: let the model act
  • When the knowledge is bigger than the context window → 4.2 RAG
  • back → handbook index

Last updated 2026-10-05

← PreviousAPI integrations: calling the model from your programNext →Tools and actions: let the model act

© 2026 Siim Liimand · SeoWeb

GitHub/AI Handbook/Tallinn, Estonia

59.4370° N, 24.7536° E — Tallinn, Estonia

↑ Top