- Home
- AI Handbook
- System Architecture
- Safety: limits and human-in-the-loop
[ System Architecture ]
Safety: limits and human-in-the-loop
Target audience: everyone | Prerequisites: 3.4 Errors and error handling
What you'll learn
After this document you will be able to:
- explain safety's main principle: the more independence you give the system, the greater the damage of an error;
- install the safeguards in four layers: what is allowed, what needs approval first, where the system stops, what leaves a trail;
- name the sensitive domains and justify why human-in-the-loop is always there;
- write safety rules into the system prompt and know why the prompt alone doesn't protect;
- run the two-question safety test before going live.
In plain terms
The question isn't “does the system err” — it does, 3.4 already showed. The question is: when it errs, how far does the error get before someone stops it? In a well-built system the error stops on a screen; in a badly built one the error reaches the customer's inbox or a bank invoice. The difference isn't in the model — it's in the design. Safety is a design question, not a hope.
Safety's main principle: independence and damage grow together
Every new independence you give the system increases the benefit and the damage at the same time. A system that only reads data and composes drafts errs safely: a wasted minute. A system that sends messages itself errs in front of customers. A system that makes payments itself errs in money — and a monetary error is often irreversible.
| System independence | An error's maximum damage | Minimum safeguard |
|---|---|---|
| reads and composes drafts | wasted time | human review before use |
| sends and changes itself | the error reaches the customer | approval points by rule + a trail |
| makes transactions itself | direct and irreversible damage | human approval always + amount limits + a stop button |
One thing doesn't change at any level: responsibility for the system's error stays with the human and the business — the customer and the law turn to you, not the model (1.6 Roles and responsibility in a project). The human's place in the system is therefore part of the design, not an afterthought.
A note in advance: if 1.2's red flags forbid the task itself, the solution isn't restrictions but the decision not to automate.
The safeguard layers: what's allowed, what needs approval, where it stops
Safety isn't one rule but safeguards (a limit built into the system's design) that work together.
Layer 1 — what the system may do
The surest protection is making the forbidden impossible, not merely forbidden. Set the system's rights up technically: what it reads, which operations it can trigger. A system that has no deletion right and no sending path can't do those things — no matter what it thinks. Everything not separately allowed is forbidden.
Layer 2 — what a human must approve first
Some actions aren't forbidden but need a signature — the more irreversible the action, the more certainly a human stands in front of it. Typical approval points:
- monetary operations — a payment, price confirmation, a discount: always a human's signature;
- first sending to a customer — the first message is the meeting trust depends on;
- changing data — correcting, deleting, or transferring a record to another system.
Approval must be fast — one criterion, one click. When reviewing takes longer than the work itself, people soon start approving blindly.
Layer 3 — what the system must stop on
A stop button (a kill switch — the ability to take the system down in one step) is needed, usable before a crisis, not during it. Three requirements:
- one step. Stopping is one click — not hunting through settings across ten menus.
- someone knows where it is. The stop button belongs to the maintainer (roles in 1.6); when they're on vacation, another person knows where the button is.
- the system also stops itself. Ten errors in a row bring the workflow to a standstill; a human reviews before continuing.
Layer 4 — the trail
Every automatic action leaves a trail (an audit trail — the record of every automatic action): what, when, and why (which rule or input triggered it). The trail is the only way, after an error, to find out what happened — and, before an error, to show that the system behaved by the rules.
In plain terms: the four layers are like a bank card: the card opens only your account (what's allowed), a large transfer needs a signature (approval), on a suspicious transaction the bank blocks the card (stop), and everything stays on the statement (trail). No single layer protects — their combination does.
Sensitive domains: where the human must always be in the chain
In most workflows approval can lighten as the system proves itself. In a sensitive domain — where one error touches health, money, or a person's rights — don't do it: there human-in-the-loop (a workflow in which a human approves the result before use) stays for the whole time, not only for exceptions.
| Domain | Why risky | minimum safety |
|---|---|---|
| Health advice | the system doesn't know the person's condition; wrong advice can harm health | answers only with catalog data; routes advice to a human; exceptions are logged |
| Monetary approvals | the error is directly and irreversibly in money | the system makes no monetary approvals on its own — all monetary operations pass a human's signature |
| Contractual decisions | one word becomes a legal obligation | the system composes the draft; a human confirms and signs |
| Hiring decisions | the decision touches a person's life; the model may reflect biases | the system sorts and summarizes; a human decides and justifies |
The pattern: the system may summarize, search, and sketch — not decide or advise. That way the risk stays with the human, who carries the responsibility.
Safety rules in the system prompt — and their limits
The system prompt is the prompt (the instruction given to the model) that works in the background of every conversation. Safety rules belong in this record, phrased on both sides:
- The prohibition clearly. “Don't give medical or drug-interaction advice.”
- The replacement just as clearly. “When a customer asks for health advice, say politely that this belongs to the pharmacist and offer the pharmacist's contact.” A prohibition without a replacement leaves the model a gap it often fills anyway.
WARNING: a rule written in the prompt isn't a guaranteed protection — over a long conversation or with a cunning input the model can get around the rule. A technical limit (the action is impossible) is stronger than a word (the action is forbidden): put the rules in the prompt, but rely on the design. How someone can override instructions — prompt injection — 3.7 Security covers.
The safety test before going live
Before the system can reach customers or real data, go through two questions in writing:
1. “What is the worst thing this system can do?”
Ask about the possible, not the probable. Concretely: “sends messages” → “sends a hundred messages with the wrong wording”; “changes data” → “deletes a record that can't be brought back”.
2. “Does the system's design prevent it?”
Require a concrete layer as the answer: “impossible, because there's no sending right”; “stops, because the approval point is in front”. If the answer is “we hope the model behaves” or “it's written in the prompt”, it isn't design — it's hope. Without a layer, the system doesn't go live.
In plain terms: like in a car: you don't ask “will I drive carefully today” but “is the seatbelt there”. Carefulness is hope, the belt is design.
Step-by-step example: Remedy House with three variants
The fictional “Remedy House” — an e-pharmacy's chat. A customer writes: “Can I take ibuprofen together with my blood pressure medication?”
Variant A — bad: the system answers itself. The system answers confidently: “Yes, you can.” Risk: the AI model doesn't know the customer's health and gives general information as firm advice; when the information is missing, it fills the gap with a plausible offer — a hallucination (the model's confidently stated but wrong answer). The error reaches the customer invisibly — and the pharmacy, which never gave the answer, is left holding the responsibility.
Variant B — limited: the prompt forbids. In the system prompt: “Don't give health advice — route to the pharmacist and give the pharmacist's contact.” The customer gets: “For that, ask the pharmacist — here's the contact and the opening hours.” The system also composes a message to the pharmacist, which a staff member answers. Considerably better — but the limit is a word: to a differently worded question (“are these two compatible?”) the prompt may no longer help.
Variant C — complete: prompt + technical limit + trail. The system technically has access only to the catalog data: it can say “Ibuprofen is in stock, €4.20, available without prescription” — and nothing else. On a health or interaction question the system routes to the pharmacist as in variant B; all routed questions are logged to the trail: what, when, where. The pharmacist's weekly summary shows which questions come up and where the routing missed.
| A — bad | B — limited | C — complete | |
|---|---|---|---|
| Limit | none | in the prompt (a word) | prompt + a technical access limit + a trail |
| To the customer | “Yes, you can” | routing to the pharmacist + a message to the pharmacist | catalog facts; the health question to the pharmacist |
| Risk | wrong advice reaches the customer | the prompt may not help | small: the facts are right, the advice stays with a human |
| If the system errs | nobody notices | the pharmacist sees the message | the trail shows what and when |
Variant C isn't “less automation” — the customer still gets stock and prices in seconds. What's safe is automated; what's sensitive stays human-in-the-loop.
Summary
- Independence and damage grow together — every new right increases the error's maximum, but responsibility always stays with the human (1.6).
- Four safeguard layers: what's allowed (a technical limit), what needs approval (a human in front), where it stops (a one-step stop button), and what leaves a trail (what, when, why).
- In sensitive domains — health, money, contracts, hiring — human-in-the-loop is always there, not only for exceptions.
- The prompt in the system message is necessary but not sufficient — the word doesn't hold, the technical limit does.
- The safety test: “What is the worst thing this system can do?” and “Does the design prevent it?” — if the answer is hope, the system isn't ready.
What's next?
- previous → 3.4 Errors and error handling
- next → 3.6 Cost management: tokens, prices, budget
- Prompt injection → 3.7 Security
- System governance and ethics → 5.7 Responsibility, ethics, and governance
- Data protection and auditability → 5.3 Security and data protection
- back → handbook index
Last updated 2026-10-05