- Home
- AI Handbook
- Fundamentals
- What is an AI model and how it “thinks”
[ Fundamentals ]
What is an AI model and how it “thinks”
Target audience: everyone | Prerequisites: none — this is the first document of the journey.
What you'll learn
After this document you will be able to:
- explain in plain language what an AI model is and what its work consists of;
- explain why the model sometimes answers convincingly but is wrong — and what to do about it;
- tell an ordinary program apart from an AI model and understand why this changes the principles of building systems;
- describe step by step how the model processes a customer inquiry about an e-invoice.
In plain terms
An AI model is like an assistant who has read an enormous library over their lifetime. When they get a question, they don't look for the answer in any file or database — they predict what text would most likely come next. That is why they are so flexible: they can answer, summarize, translate, write. But behind that same trait lies a weakness: they don't know, they guess — and they can guess convincingly, stating something untrue in a confident voice.
Keep that in mind and the whole handbook becomes easier to understand: we are not building a system that executes commands exactly, but a system that makes a very good offer — and our job is to check it before use.
How an AI model actually works
Input → prediction → output
In plain terms: the model takes in text and searches for the most likely next token to continue it — many times over. That is the whole of its “thinking.”
For the technical reader: a large language model (LLM) is a neural network (a computational model whose parameters — the so-called weights — are billions of numbers) that at every step computes the probability of every possible next token based on the context. The answer is born piece by piece: the model reads what has been written so far, picks the next piece, adds it to the context, and repeats the loop. Your text is the input; the answer assembled piece by piece is the output.
An important conclusion from this: the model has no database to pull answers from. It does not look up your e-invoice or read the customer registry — if you want it to work with your data, you must put the data into its input yourself.
Training vs inference
In plain terms: training is a years-long education — reading a huge amount of text and learning how one text continues after another. Inference is what you do every day — you ask and get an answer. The model learns nothing from your conversation: tomorrow it starts again with the same knowledge.
During training, the model is shown enormous amounts of text and its weights are tuned so that the prediction becomes more accurate. Your part is inference (the model's actual run): you send an input and get an answer.
An important consequence: the model's knowledge is “frozen” at the moment of training — it knows nothing about events after training and does not learn from your data on its own. When needed, up-to-date information must be brought into the context.
The token — the model's piece of a word
In plain terms: the model does not read letters or whole words — it cuts the text into pieces called tokens (a token is a piece of text somewhere between a few characters and a whole word). To the model, a token is one “word.”
Example. In the model's eyes the sentence “Invoice no. 2026-14 is unpaid” may split like this:
["Invoice", " no", ".", " 2026", "-", "14", " is", " un", "paid"]
Note that longer words get broken into pieces (“unpaid” → “un” + “paid”). Text in languages other than English — Estonian included — therefore often consumes more tokens than English, and many services charge by the token.
The context window — how much the model “remembers” at once
In plain terms: the context window (the amount of text the model can see at once) is like a desktop surface: whatever doesn't fit on the surface stays invisible, even if it exists elsewhere in a drawer.
For comparison: a book page holds about 500 words, but some of today's models fit tens of thousands to millions of tokens into the window — dozens of books at once. Still not the whole world. Two practical conclusions:
- If the model “forgets” something from the beginning of a conversation, it is not willful — the information is outside the window, or the model's attention (the mechanism that distributes weight across the parts of the input) has been captured by other information.
- Every input and output consumes context — a long automated work cycle fills the window quickly.
Large and small models — why the choice matters
In plain terms: a large model is an experienced specialist — capable, but slow and expensive. A small model is a diligent intern — fast and cheap, but it needs more precise instructions and makes more mistakes on demanding tasks.
Large models can handle more complex logic and long instructions, but they cost more and answer more slowly. Smaller ones suit simple, high-volume jobs — for example classifying mail (“invoice,” “complaint,” “order”). The system builder's rule: start smaller and move up to a larger model only where the smaller one cannot cope in testing.
Why the model is not always right (hallucination)
In plain terms: the model is an excellent persuader. It can state something made up in a confident voice and as a well-constructed text — this is called a hallucination (an answer the model states confidently but that is false).
Why does this happen? Because the model predicts the next token based on probability, the answer that sounds right and the answer that is right are, from its point of view, similar tasks. If a fact is weak or missing from its knowledge, it does not say “I don't know” — it continues with the most likely text. So it may create, for example, an email address that sounds plausible but belongs to no one.
The classic business example: ask for sales figures without including any figures — the model may offer rounded, plausible sums that match no actual data.
Three practical rules for reducing the risk:
- Give the model the facts. Don't rely on the model's memory — put the necessary information (a document, a table, database fields) directly into the input. Then the answer is based on your data, not statistical memory.
- Demand a source. For critical information, ask the model to rely on existing text and show where the information comes from. The system must also allow the answer “no source found.”
- Critical decisions stay with a human. Money, contracts, legal positions — a human reviews before anything leaves the system: the higher the price of an error, the stronger the checkpoint.
The difference between a program and an AI model
| Ordinary program | AI model |
|---|---|
| Deterministic — the same input always yields the same result | Statistical — the same input can yield a different answer on different runs |
| Works by rules written by a programmer | Has learned from examples — the rules are hidden in internal numbers |
| The error is predictable: the program crashes (stops working) or gives an error message | The error is convincing: a wrong answer looks like the right one |
| Cannot do anything it wasn't programmed to do | Can generalize to situations it was never shown |
From this follows the most important thing: the principles of building an AI system are different — in an AI-based system, randomness is a natural part, not an error. Therefore:
- we build in automated checks (e.g., “does the answer contain the invoice number, and does it exist in the database?”) that catch bad answers;
- we place human review where the automated check cannot cope;
- we design a fallback (a prepared alternative the system uses when the primary path fails): a wrong answer is not an exception but a situation the system must be able to handle.
This viewpoint — the model as an uncertain but capable component surrounded by an orderly system — is the core of the whole handbook.
A step-by-step example: how the model processes a customer inquiry
Suppose a customer writes: “Hello! Has my e-invoice no. 2026-14 been paid yet?”
- The customer's letter reaches the system. The system (not the model) reads the letter and decides that it must be answered.
- The text is split into tokens. The letter breaks into tokens and is converted into numbers the model can process.
- The system assembles the context. A crucial step: the letter is augmented with the customer's data (e.g., “invoice 2026-14: paid 12.03.2026”) and instructions (“reply briefly, reference the invoice number”). Now the input contains everything a correct answer needs.
- The model predicts the answer piece by piece. The model reads the entire context and writes the answer one token at a time.
- The answer reaches the customer. For example: “Hello! Yes, invoice no. 2026-14 has been paid — the payment arrived on 12.03.2026.”
- Where it can go wrong and what the human sees. If step 3 provided no data, the model may hallucinate — invent a payment date that sounds plausible but is not true. That is why the system shows the human where the data came from and what the answer is based on; an answer with a wrong number is routed to a human for review.
Note: the model's role is only step 4 — everything else (fetching data, assembling the context, checking the answer) is ordinary, reliable software. This division of labor (called orchestration — a system that coordinates the cooperation of the parts) is the key to a well-functioning AI system.
Summary
- An AI model does not look up answers; it predicts the next token — that is why it is flexible, but right without a guarantee.
- The token is the model's unit of work — the piece of text the input is cut into; the number of tokens affects the price and how much fits into the context window.
- Hallucination is not junk but a natural consequence of prediction logic — supplying facts, demanding a source, and a human checkpoint help against it.
- An ordinary program crashes when it errs; an AI model gives a convincing wrong answer — that is why the model must be surrounded by a system that checks, limits, and when necessary routes to a human.
What's next?
- Next document → 1.2 Capabilities and limits: what is worth automating
- back → handbook index
Last updated 2026-10-05