- Home
- AI Handbook
- Practice
- Inputs and data preparation
[ Practice ]
Inputs and data preparation
Target audience: everyone | Prerequisites: 2.3 Basic workflows
What you'll learn
After this document you will be able to:
- name the four properties that distinguish a good input dataset from a bad one;
- recognize the five most typical data problems and fix them before handing the data to the model;
- format the data into a uniform structure with clear labels;
- decide how much data to put into the request and stick to the privacy principle already during preparation;
- walk the whole process through with one example asset: call transcriptions (a written record of a call) as structured input.
In plain terms
The same AI model gives a reliable answer with proper input data — and the same model, with a messy, faulty dataset, gives a poor answer or invents the missing parts itself. When the answers are bad, look first at what went into the input. Three steps — cleaning, formatting, and a sensible, safe amount — are done before the model ever gets to deal with the thing.
What good input is
From document 1.3 you know that the prompt (the instruction given to the model) is the entire content the model sees at once — including all the input data. Weak data gives just as unpredictable a result as weak instructions. Four properties to check on every part of the input:
- Relevant. It contains what this task needs — and nothing else. Leftover information isn't a harmless addition: it is noise that makes it harder for the model to find the important part, and every extra word spends tokens (a token — the word-sized piece the model cuts text into).
- Complete. All the facts a correct answer needs are in the input. Missing information doesn't vanish anywhere — the model fills the gaps with general knowledge or invents something extra (a hallucination — a confidently stated answer from the model that isn't based on the facts; see 1.1).
- Uniformly formatted. The same thing is written everywhere in the same way: dates in one form, numbers with one separator, a name always identical. Different spellings force the model to infer whether “Katelin” and “Katlin” are the same person.
- Timely. The data is up to date: this year's price list, not last year's. Outdated data is worse than missing data — an answer based on it looks right, but isn't.
Cleaning: five typical problems
In plain terms: cleaning (fixing data errors before handing the data to the model) is like washing the dishes before cooking — nobody wants to do it, but skipping it shows in the taste of every dish. The five problems below cover most real-world data errors.
| Problem | Typical example | Fix before handing to the model |
|---|---|---|
| Duplicates and same-content records | The same client twice in the database under different names; two transcriptions of the same call | Merge the duplicates, keep the more current version — duplicates bloat the input and create contradictions |
| Half-finished records | An order without a quantity; a contact without a phone number | Fill in the missing piece from a reliable source or mark it explicitly “information missing” — don't leave a gap for the model to guess |
| Different date and number formats | “2.10”, “2026-10-02”, “Monday”; prices both as “2,50” and “2.5” | Normalize to one form (e.g., dates as 2026-10-02) and numbers with the same separator already before putting them into the input |
| Transcription errors | Speech recognition writes the name now “Katlin”, then “Katelin”, then “Katlynn” | Keep a correction table for names and products and replace every variant with the correct spelling |
| A mix of languages | A letter in Estonian where the quantities and product names are in English | Feed the input in one language as much as possible, or state in the instructions that the input may be mixed-language |
Cleaning doesn't have to be a thorough renovation — catching these five errors before they go into the input already raises the quality noticeably, because fixing the answer after is always more expensive than fixing the input before.
How to format data for the model
In plain terms: formatting (giving a uniform shape) means that the data is arranged the way it will be read — not the way it was collected. You find the answer in an organized table immediately; from the pages of a notebook, almost never.
Three rules:
- Uniform structure. Every record with the same fields and in the same order — client, order, date, status. Then the model knows to look for exactly the same thing at every record.
- Clearly separated sections. Use labels (a label — a clearly separated section of the input, e.g. [QUESTION]) that build rooms with headings into the input: here is the question, here the data, here the history.
- Data as a list, without conversational wrapping. “quantity: 2 pcs” — not “so she then said she'd take maybe two pcs, let's say”.
The same information in two forms:
Unorganized input:
So this Katlin Tamm asked where her order 312 has gotten to, she ordered
2 pcs of that cleaning agent and the order was placed on September 28,
last time around she got a 10% discount.
The same information, organized:
[CUSTOMER_QUESTION] Where is my order no. 312?
[ORDER]
- client: Katlin Tamm
- contents: 2 pcs of cleaning agent
- submitted: 2026-09-28
[PREVIOUS_RESOLUTION] 3-day delay → 10% discount
The content is identical, but in the organized version the model knows exactly what is the question, what the data, and what the history. Even more important: every following record can fill this form — nobody can repeat a human's freely written paragraph.
How much data to give
In plain terms: the context window is like a travel suitcase — it holds a lot, but not the whole household, and the more you squeeze in, the harder it is to find the more important things. Give everything the model needs and nothing else.
- The context window (the amount of text the model can see at once) is limited. Every input spends tokens — that means money and time, too; text in Estonian more than text in English (see 1.1).
- Prioritize. Ask of every dataset: “Is this dataset truly necessary for the answer?” A monthly summary for the recipient is often exactly as good when it is based on the five most important records, not fifty — only cheaper and more precise.
- The order matters. Put the most important at the beginning and the end: in a long input, the middle part gets the least attention. This is one of 1.3's five most common mistakes — important information buried in the long middle (see 1.3 Prompt fundamentals) — and it holds even when the input is filled with data.
- If there is too much data to fit into the input at all, the solution is not “squeeze it in anyway” but retrieval: the system pulls out only the relevant passages based on the question and gives them to the model — this is called RAG (retrieval-augmented generation — a system that looks up the answers in the company's own documents; see 4.2).
Privacy already during preparation
In plain terms: data you have never sent out cannot ever leak either. The cheapest data protection is not needing to do it — achieved already when assembling the input, not with the system's security afterwards.
About every field ask: does the task truly need it? Often not. The client can be “Client #142” in the input — that is fully enough for the model, because it has to tell records apart, not know a person. The same goes for personal ID codes, addresses, and other personal data: replace or remove them before sending, not after.
When the input data flows into the system automatically — every call and every letter really goes into the input — nobody manages to review them by hand. Then cleaning and anonymization must be an automated step of the system; how to place such steps into a workflow is told by 2.3. Security as a whole — keys, data, malicious instructions — is covered by 3.7 Security.
A step-by-step example: CallHelp's call transcriptions
The situation: the phone support of “CallHelp”. Over every night, 15–30 client calls pile up that an employee cannot manage to listen through in the morning. The system transcribes the calls, but the texts are so ugly that giving them straight to the model makes the answers a lottery.
Step 1 — the initial chaos. One real transcription:
hi, i ordered something but there's no sign of it, no. three-one-two.
I'm Katelin, asking for the second time already, the first time you said
it would come tomorrow, that was on a Monday... otherwise last time
it got solved when you gave 10% off. voice: Katlynn Tamme
Step 2 — identify three problems (from the cleaning table above):
- Transcription errors — the client's name written in two different ways (“Katelin”, “Katlynn”) and the order number in words (“three-one-two”);
- Different formats — the date only in words (“Monday”) — when was it?
- Half-finished data and conversational wrapping — “i ordered something”: which product, how much, when?
Step 3 — clean. The employee builds a name correction table and fills the gaps from the system's data:
| Spelling identified in the call | Correct spelling |
|---|---|
| Katelin, Katlynn | Katlin Tamm |
Dates normalized: “Monday” → 2026-09-28 (from the order data). The number from words into digits: “three-one-two” → 312. The half-finished data from the order record: 2 pcs of cleaning agent.
Step 4 — formatted as structured input:
[ROLE] You are the CallHelp customer service assistant — businesslike and friendly.
[CUSTOMER_NAME] Katlin Tamm (Client #142)
[CUSTOMER_QUESTION] Where is my order no. 312? The client is asking for the second time.
[ORDER]
- number: 312
- submitted: 2026-09-28
- contents: 2 pcs of cleaning agent
- status: on the way, estimated 2026-10-06
[PREVIOUS_RESOLUTION] 2026-09-15: 3-day delay → 10% discount
[TASK] Draft a reply to the client of up to 80 words, in Estonian.
Step 5 — the result. Answers based directly on the raw transcript were a lottery: one run the model offered a product, one run misspelled the name, one run guessed the date. With cleaned and structured input, all the facts come from the data, not from guesswork — the answers are more reliable and repeatable. Because the steps are the same every morning, the system can run this sequence on its own; how to build it as a workflow is taught by 2.3.
Summary
- Good input is relevant, complete, uniformly formatted, and timely — all four are checkable before handing the data to the model.
- Cleaning covers five typical problems: duplicates, half-finished records, different formats, transcription errors, and a mix of languages — every one of them is simple to fix if done before, not after.
- Formatting makes the work repeatable: a uniform structure, clear labels, and a clean list without conversational wrapping.
- The context window is limited and attention is uneven — give only what is needed, keep the important at the beginning and the end, and for large datasets use retrieval (RAG).
- Privacy starts with preparation: less data means less risk — “Client #142” is often enough.
What's next?
- previous → 2.3 Basic workflows
- next → 2.5 Your first end-to-end automated workflow
- When there is too much data → 4.2 RAG: using your own data
- back → handbook index
Last updated 2026-10-05