- Home
- AI Handbook
- System Architecture
- Errors and error handling
[ System Architecture ]
Errors and error handling
Target audience: technical + non-technical readers who want to go deeper | Prerequisites: 3.3 Tools and actions
What you'll learn
After this document you will be able to:
- name the five error types and each one's usual solution;
- put a limit on a retry (trying again) and explain why endless repetition makes things worse;
- plan a fallback (the designated B path) that takes over work the main path couldn't finish;
- decide what to write down about errors (logging — the recording of events);
- say when a human is alerted and when a single entry in the log is enough.
In plain terms
Errors are not a system breakdown but part of its ordinary day. Just as a store's cash register can jam and the salesperson takes the money by hand, the flow has pre-written: what happens when the AI model doesn't answer in time, when the answer comes in the wrong form, or when the email service is down. If your paths are ready, an error stays an entry in the log — not an emergency.
This document's principle is: errors are normal and planned for — they are not an emergency. 2.5 named error types and promised that every error has an address — now we look at what to do about them.
Error types: what can go wrong
Most production errors of workflows fall into one of five types:
| Error type | What happens | Usual solution |
|---|---|---|
| Timeout | the model doesn't answer in time — the request is left waiting for a response | retry a limited number of times; if that doesn't help, the fallback |
| Rate limit exceeded | too many requests in a short time — the service pushes them back | retry with a wait; slowing the requests down — the technical side in 3.1 |
| Broken output | the answer has drifted off the format — the JSON (the structured data format) can't be parsed because the model added a chatty sentence | the automatic check finds there are no fields — entry to the fallback path (2.2's check) |
| Content error | the form is right but the content wrong: a hallucination (the model's confidently stated but wrong answer; 1.1) or a wrong database answer | prevention (facts from data, not the model's memory) and a checkpoint before the output (1.5's fifth part) |
| External error | another service isn't working: the email service is down, the database doesn't respond | saving the work and continuing when the service recovers |
Timeout and rate limit are temporary — the same request may go through a moment later. Broken output and content error have already happened — a new answer will err the same way. An external error isn't this flow's own — it belongs to another machine.
What does broken output look like? The system expects a response like this (shown shortened):
{
"type": "return",
"order_number": "1187"
}
If the model answers instead: “Sure! The message's type is a return…” — the check finds no fields at all: the content is almost right, but the form is not. That's why the check looks at the form first and only then the content.
In plain terms: a timeout and a rate limit are a doorbell after which someone shouts “wait, I'm coming!” — ring again a moment later and everything goes through. A broken output and a content error have already happened — the check and the fallback fix them, not another try.
Retry: when and how many times
Retry is the simplest means: the same request is sent again. It suits only temporary errors — timeouts and rate limits. The main rule:
For a temporary error, retry with a stopping time — for example three attempts and a wait between them — not forever.
Why rapid-fire retrying makes things worse:
- Retry adds load to load. If a service is slow, all the services retrying every second keep it down much longer — the retry becomes the cause of the delay.
- With a rate limit, the retry becomes the error itself. Without a wait, every new request is exactly what the limit forbids: the more you try, the more gets pushed back.
- Endless retrying keeps the work hanging. The message stays a half-finished job, nobody knows where it is, and the customer waits.
That's why the stopping time is part of the system: for example three attempts, each waiting longer than the previous; when the limit is used up, there's no fourth try — the fallback is taken. (Setting a timeout on the request — an API call detail, see 3.1.) A retry is also a new model call and new money — the cost side, 3.6 covers.
In plain terms: a retry is the second and third press of the doorbell — polite and often enough. Pressing every second doesn't open the door faster, but wakes up the neighbor — that's why the number of attempts and the wait are written down in advance.
The fallback: the planned path when the main path doesn't work
The fallback is the turn to a human familiar from 1.5's anatomy, but it isn't the only form — ordinary flows use three, often side by side:
- A rule-based standard answer. When the real answer can't be composed with certainty, the system sends a pre-written, tried-out text: “Thank you for writing — we'll get back to you within one day.” — directly or through a human's approval. The promise is fulfillable, because the message is already in the human's list. The model may have erred, but the error didn't reach the customer.
- Saving the work and manual continuation. When the external part is temporarily down, the system saves the unfinished work into a queue and continues when the service recovers — the work isn't lost, it moves later.
- Routing to a human. When no automatic path is trustworthy, the work goes into the human review list — just as the “other” inquiries went to Piret in 2.5. This isn't a failure but part of the design (1.5) — an uncertain estimate doesn't have to decide by itself.
Behind these stands one principle: a fallback is always better than a wrong answer to the customer. A customer who received a correct “we'll get back to you” message stays satisfied; a customer who received a confident-toned but wrong answer loses trust — three correct messages later won't restore it. What the system may undertake on its own and what needs a human's approval (human-in-the-loop) is a safety decision — see 3.5.
In plain terms: a fallback is ordering a taxi in advance before the first one breaks down: with the main path jammed, the work moves on the planned parallel path, not in a random direction. The customer is left with the impression that everyone had a plan — because they did.
Logging errors: what to record and why
When the error is solved, for the customer the matter is done — for the developer it's only beginning. By 1.5, every error has an address, and only an entry shows the address: without a log, what happened remains “something went wrong”; with a log, there's the place to go to.
What to record about every solved error:
| Entry part | Example | What for |
|---|---|---|
| time | October 13, 2:32 PM | do errors pile up at a particular time (peak hour)? |
| input | the message's type and the order number | which input led to the error? |
| error | a timeout after 20 seconds | which type of error was it? |
| path taken | retry — the second attempt succeeded | was the main path, a retry, or the fallback enough? |
A simple rule for alerting a human: alert when the fallback fires, and when the same error repeats a set number of times (for example five times a day). One retry needs no human; a fallback and a repeating error mean something has changed — someone has to know. In more depth, 4.6 Monitoring in production covers.
In plain terms: the log is a driving diary: three timeouts on Wednesday, all solved with a retry, are three entries — not three crises. When the diary shows that timeouts pile up around midday, you have a fact to decide on, not a feeling.
Step-by-step example: three incidents in the e-store
2.5 built the HomeCraft Store's return-request flow and tracked the first week with three numbers. The same week's three incidents:
| # | Incident | What happened | Which path the system took | Result for the customer |
|---|---|---|---|---|
| a | timeout | the model answered in 30 seconds — the request exceeded the 20-second wait time | retry: the system tried again and the second time the answer came within seconds | the customer's message wasn't lost: the draft reached Piret for approval a minute late — for the customer an ordinary confirmation chain, just a small delay |
| b | broken output | the model answered in free text (“Sure! The message's type is a return…”) — the automatic rule found no fields | fallback: the check caught the breakage, the system took the standard answer and put the message on Piret's list | the customer immediately got a correct answer, the real answer within the promised day |
| c | external error | the email service was down for 10 minutes — the reply messages didn't go out | saving the work: the records waited in the queue and went out when the service recovered | the customer got the answer about 10 minutes later — knows nothing about it |
Three errors — and none reached the customer as an error. The work left was three entries in the log and one decision: the timeout's wait time was extended, because the timeouts piled onto the same half-day — the decision was made from the diary during working hours, not at night.
In plain terms: the first week gave three errors and zero emergencies — planned-for errors stay entries in a log and sometimes one small decision, not a customer's problem or a team's night.
Summary
- Five error types, each with a usual solution — nothing needs to be invented in the middle of a crisis.
- Retry only for temporary errors and with a stopping time (for example three attempts, the gap lengthening): rapid-fire retrying adds load to overload and becomes an error itself.
- The fallback is the designated B path — a standard answer, saving the work, or a turn to a human — and always better than a wrong answer to the customer.
- The log turns an error into a lesson: time, input, error, and path taken; a human is alerted when the fallback fires and on a repeating error.
- Errors are normal and planned for — not an emergency: a system with paths ready lives through every error as one entry.
What's next?
- previous → 3.3 Tools and actions
- next → 3.5 Safety: limits and human-in-the-loop
- Monitoring in production → 4.6
- back → handbook index
Last updated 2026-10-05