SeoWeb
  • Work
  • Services
  • CV
  • Contact
    • AI
  1. Home/
  2. AI Handbook/
  3. Agents & Evaluation/
  4. RAG: using your own data as a source of answers

[ Agents & Evaluation ]

RAG: using your own data as a source of answers

Target audience: technical + deep-diving non-technical | Prerequisites: 4.1 Agent systems, 3.2 Context management

What you'll learn

After this document you will be able to:

  • name the three paths for getting your knowledge to the model;
  • describe the RAG flow (retrieval-augmented generation) in four steps: preparation, retrieval, context, answer with a source;
  • explain in non-technical terms how retrieval (finding the piece that matches the question) works through similarity;
  • say why the source (where the answer came from) is a mandatory part of the answer;
  • maintain a knowledge base (a collection of the company's own documents) as a living system and recognize when RAG is not needed.

In plain terms

The model learned its knowledge from the world, not from your office — internal rules, the price list, and the billing guide are foreign to it. The core idea of RAG: don't try to teach the model things by heart; instead, hand it the right page before every answer. The system finds the piece of the documents that matches the question, puts it in front of the model together with the question, and the answer states where it came from.

Three paths when the model must know your business

When a model must answer questions about your company's affairs — internal rules, the price list, rule files — there are three paths:

1. Put everything in the prompt. 2.4 showed how to pass data along with the request — and when the knowledge is small and unchanging, this is the right path. But the limit arrives: the context window (the amount of text the model sees at once) fills up. With a large body of knowledge it doesn't pay to put it in the prompt: a 200-page manual is ~150,000 tokens, which gets billed again with every question (see 3.6), and in a large volume of information much of it would go unheeded (see 3.2) — even if a large window could technically hold it.

2. Retrain the model. Expensive, time-consuming, and usually pointless: facts change, and training has to be redone after every update — and the answer has no source. Training suits teaching style and form, not storing changing facts.

3. RAG — the third path. The knowledge stays in your documents; for every question the right piece is retrieved from there, placed into the context, and the model answers based on it. The model isn't retrained — only the document collection is updated.

In plain terms: the first two paths are “put everything on the table” and “memorize it” — the first fills up and gets expensive with a large body of knowledge, the second goes stale the moment a price changes. The third is an archive and an assistant who finds the currently needed page in it.

When is the third path not needed? Three cases: the knowledge is small and stable — then everything fits in the prompt (see 2.4); the answer must be a one-hundred-percent exact computation — then function calling helps, not RAG (see 3.3); the data is confidential — then the solution is not RAG but reducing the data and choosing a provider carefully (see 3.7 Security).

The RAG flow in four steps

1. Document preparation: chunking

The first step doesn't involve the model at all: documents are cut into searchable pieces — chunking. Two rules:

  • One topic per piece. Cutting follows the headings: a chapter or a subsection together with its heading. A too-large piece brings a pile of noise; a too-small one becomes meaningless — a passage “you may deduct 40 percent” without a heading doesn't say which expense it's about.
  • Uniform structure. Each piece knows which document and chapter it belongs to and when it is valid — it is exactly this information that later reaches the answer as the source.

2. Retrieval: how the right piece is found

The answer usually lives in one piece among thousands — how does the system find it? Through similarity.

Each piece is represented numerically — as a vector (embedding: a numerical representation of text that lets you find similar texts). Imagine that each piece gets a fixed shelf location in a library and texts with similar content end up side by side on the shelf. The question is placed into the same kind of representation and the retrieval brings its nearest neighbors. The question and the piece may differ in wording, but in meaning they remain side by side on the shelf.

If retrieval doesn't find the right piece — coarse pieces, a vague question — the model is left without facts and starts guessing; 3.4 Errors and error handling covers such cases.

3. Placing the pieces into context together with the question

In the context window, the system assembles an instruction (“answer only based on the source pieces; if the answer isn't here, say so”), the 2–5 best pieces found, and the question. From then on the situation is the same as with any ordinary request (see 3.2) — except that the knowledge piece arrives in the window from retrieval, not from you.

4. The answer with a source

The model formulates the answer and shows where it came from. As structured output (see 2.2) it looks roughly like this:

{
  "question": "May remote-work communication costs be deducted from payroll taxes?",
  "answer": "Yes, if the remote-work agreement is in writing and the cost is documented.",
  "sources": [
    { "document": "Billing Rules Handbook", "chunk": "4.2", "valid_from": "2026-09-01" }
  ]
}

The whole flow in one picture:

   A 200-page rules handbook
        │
        ▼
   1. CHUNKING: ~600 pieces → each piece into a numerical representation (a “shelf location”)
        │
        ▼
   QUESTION: “may communication costs for remote work be deducted?”
        │
        ▼
   2. RETRIEVAL: the question into the same representation → nearest neighbors (chunk 4.2, chunk 2.3)
        │
        ▼
   3. CONTEXT: instruction + chunk 4.2 + chunk 2.3 + question
        │
        ▼
   4. ANSWER + SOURCE: “Yes, if ... — the handbook, section 4.2”

The source is mandatory: an answer you can check

The source is not a decoration — three practical reasons why an answer without a source must not be allowed:

  1. Trust. The user sees where the answer came from and doesn't have to believe blindly.
  2. Verification. The human in the loop (a workflow where a human approves the result before it is used) can open section 4.2 and compare — the assessment becomes a checkable claim.
  3. Hallucination detection. A hallucination (the model's confidently stated but wrong answer, see 1.1) doesn't disappear in RAG: if the right piece didn't come from retrieval, the model may still make something up convincingly. The source makes this visible: if sources are missing or don't support the statement, it's a guess. That's why the instruction also contains the rule “if the pieces don't contain the answer, say so”.

In plain terms: an answer without a source is like a quote without page numbers — trust it only when no checking is needed. With money and rules, checking is always needed.

RAG is a living system: maintenance

Documents change: a price is updated, a rule is revoked, a condition is added. A RAG system doesn't automatically follow along — and this is exactly where the most dangerous error lives:

A stale piece is worse than a missing one. A missing answer brings an honest “not found” and the human searches themselves; a stale piece brings a confident, sourced old answer — and nobody checks, because the source looks trustworthy.

To prevent this, three maintenance rules:

  • Update = old out, new in. When a rule changes, the old piece must be removed from the knowledge base — not left next to the new one, otherwise two answers live in the base and the system picks between them unpredictably.
  • One truth per topic. One valid version per document, a validity date attached to every piece.
  • Regular checks. After every document change, ask the system a few known questions and see whether the sources are up to date. Whether the system answers correctly — how to measure that is the topic of evaluation (see 4.5 Evaluation).

RAG is not a one-time build but a living system: maintaining the knowledge base is maintaining the system.

A step-by-step example: the Number Bureau rules knowledge base

The fictional Number Bureau (see 1.6) — owner Anu, five accountants, and Jaak, who builds the systems (and whose hand produced the invoice-status tool in 3.3). The next recurring annoyance: the accountants ask about billing rules — “may phone and internet costs still be deducted from payroll taxes when the person also works from home?” — and someone has to dig the answer out of a 200-page internal file.

Putting this file into the window isn't worth it: ~150,000 tokens billed again with every question (see 3.6) and a large share of so much information would go unheeded (see 3.2) — even if a large window could technically hold it. Nobody is going to retrain the model. Jaak builds RAG:

1. Preparation. The handbook is cut by chapters into about 600 pieces; each piece with its heading (e.g. “4.2 Costs that may be deducted from payroll taxes”), the document name, and a validity date.

2. Retrieval. Each piece gets a numerical representation; the accountant's question is placed into the same representation and retrieval brings the two nearest pieces: “4.2 Deductible costs” and “2.3 Remote-work conditions”.

3. Context. The instruction (“answer only based on the pieces; if there is no answer, say so”), the two pieces, the question.

4. The answer with a source:

“Yes — for remote work, communication costs are deductible if the agreement is in writing and the costs are documented. (Source: Billing Rules Handbook, section 4.2, valid from 2026-09-01.)”

Before: every question took ten minutes of digging through folders. After: seconds — and the accountant verifies the answer against the source before telling it to the client.

And then the inevitable happens. On October 1 the rule changes. Jaak puts the updated file into the base, but one piece of the old version stays behind. An accountant asks the same question and gets the answer according to the old rule — confidently and with a source. Solution: old piece out, new one in, validity dates distinguish them — a one-time fix becomes a fixed update process.

In plain terms: building took Jaak a week; maintenance takes five minutes with every change. The system is good exactly as long as someone keeps the documents up to date — that “someone” must be named before the first question.

Summary

  • Three paths: everything in the prompt, retraining (expensive, pointless for changing facts), RAG (retrieve the right piece and put it into context) — for large and changing knowledge, the third is the only one that works.
  • The flow in four steps: chunking (one topic per piece, uniform structure) → retrieval (through similarity, via vectors) → placing the pieces into context with the question → the answer with a source.
  • The source is mandatory: trust, human-in-the-loop verification, and hallucination detection.
  • RAG is a living system: on update the old piece must go out, because a stale piece gives a confident wrong answer.
  • When RAG is not needed: the knowledge fits in the prompt; the answer must be a one-hundred-percent exact computation (function calling, 3.3); the data is confidential (reduce the data, 3.7).

What's next?

  • If the knowledge is larger than it makes sense to put in the window → 3.2 Context management
  • If retrieval doesn't find the right piece → 3.4 Errors and error handling
  • Whether RAG answers correctly → 4.5 Evaluation
  • previous → 4.1 Agent systems
  • next → 4.3 Long-term memory and state management
  • back → handbook index

Last updated 2026-10-05

← PreviousAgent systems: what they are and when you need oneNext →Long-term memory and state management

© 2026 Siim Liimand · SeoWeb

GitHub/AI Handbook/Tallinn, Estonia

59.4370° N, 24.7536° E — Tallinn, Estonia

↑ Top