SeoWeb
  • Work
  • Services
  • CV
  • Contact
    • AI
  1. Home/
  2. AI Handbook/
  3. System Architecture/
  4. Errors and error handling

[ System Architecture ]

Errors and error handling

Target audience: technical + non-technical readers who want to go deeper | Prerequisites: 3.3 Tools and actions

What you'll learn

After this document you will be able to:

  • name the five error types and each one's usual solution;
  • put a limit on a retry (trying again) and explain why endless repetition makes things worse;
  • plan a fallback (the designated B path) that takes over work the main path couldn't finish;
  • decide what to write down about errors (logging — the recording of events);
  • say when a human is alerted and when a single entry in the log is enough.

In plain terms

Errors are not a system breakdown but part of its ordinary day. Just as a store's cash register can jam and the salesperson takes the money by hand, the flow has pre-written: what happens when the AI model doesn't answer in time, when the answer comes in the wrong form, or when the email service is down. If your paths are ready, an error stays an entry in the log — not an emergency.

This document's principle is: errors are normal and planned for — they are not an emergency. 2.5 named error types and promised that every error has an address — now we look at what to do about them.

Error types: what can go wrong

Most production errors of workflows fall into one of five types:

Error typeWhat happensUsual solution
Timeoutthe model doesn't answer in time — the request is left waiting for a responseretry a limited number of times; if that doesn't help, the fallback
Rate limit exceededtoo many requests in a short time — the service pushes them backretry with a wait; slowing the requests down — the technical side in 3.1
Broken outputthe answer has drifted off the format — the JSON (the structured data format) can't be parsed because the model added a chatty sentencethe automatic check finds there are no fields — entry to the fallback path (2.2's check)
Content errorthe form is right but the content wrong: a hallucination (the model's confidently stated but wrong answer; 1.1) or a wrong database answerprevention (facts from data, not the model's memory) and a checkpoint before the output (1.5's fifth part)
External erroranother service isn't working: the email service is down, the database doesn't respondsaving the work and continuing when the service recovers

Timeout and rate limit are temporary — the same request may go through a moment later. Broken output and content error have already happened — a new answer will err the same way. An external error isn't this flow's own — it belongs to another machine.

What does broken output look like? The system expects a response like this (shown shortened):

{
  "type": "return",
  "order_number": "1187"
}

If the model answers instead: “Sure! The message's type is a return…” — the check finds no fields at all: the content is almost right, but the form is not. That's why the check looks at the form first and only then the content.

In plain terms: a timeout and a rate limit are a doorbell after which someone shouts “wait, I'm coming!” — ring again a moment later and everything goes through. A broken output and a content error have already happened — the check and the fallback fix them, not another try.

Retry: when and how many times

Retry is the simplest means: the same request is sent again. It suits only temporary errors — timeouts and rate limits. The main rule:

For a temporary error, retry with a stopping time — for example three attempts and a wait between them — not forever.

Why rapid-fire retrying makes things worse:

  1. Retry adds load to load. If a service is slow, all the services retrying every second keep it down much longer — the retry becomes the cause of the delay.
  2. With a rate limit, the retry becomes the error itself. Without a wait, every new request is exactly what the limit forbids: the more you try, the more gets pushed back.
  3. Endless retrying keeps the work hanging. The message stays a half-finished job, nobody knows where it is, and the customer waits.

That's why the stopping time is part of the system: for example three attempts, each waiting longer than the previous; when the limit is used up, there's no fourth try — the fallback is taken. (Setting a timeout on the request — an API call detail, see 3.1.) A retry is also a new model call and new money — the cost side, 3.6 covers.

In plain terms: a retry is the second and third press of the doorbell — polite and often enough. Pressing every second doesn't open the door faster, but wakes up the neighbor — that's why the number of attempts and the wait are written down in advance.

The fallback: the planned path when the main path doesn't work

The fallback is the turn to a human familiar from 1.5's anatomy, but it isn't the only form — ordinary flows use three, often side by side:

  1. A rule-based standard answer. When the real answer can't be composed with certainty, the system sends a pre-written, tried-out text: “Thank you for writing — we'll get back to you within one day.” — directly or through a human's approval. The promise is fulfillable, because the message is already in the human's list. The model may have erred, but the error didn't reach the customer.
  2. Saving the work and manual continuation. When the external part is temporarily down, the system saves the unfinished work into a queue and continues when the service recovers — the work isn't lost, it moves later.
  3. Routing to a human. When no automatic path is trustworthy, the work goes into the human review list — just as the “other” inquiries went to Piret in 2.5. This isn't a failure but part of the design (1.5) — an uncertain estimate doesn't have to decide by itself.

Behind these stands one principle: a fallback is always better than a wrong answer to the customer. A customer who received a correct “we'll get back to you” message stays satisfied; a customer who received a confident-toned but wrong answer loses trust — three correct messages later won't restore it. What the system may undertake on its own and what needs a human's approval (human-in-the-loop) is a safety decision — see 3.5.

In plain terms: a fallback is ordering a taxi in advance before the first one breaks down: with the main path jammed, the work moves on the planned parallel path, not in a random direction. The customer is left with the impression that everyone had a plan — because they did.

Logging errors: what to record and why

When the error is solved, for the customer the matter is done — for the developer it's only beginning. By 1.5, every error has an address, and only an entry shows the address: without a log, what happened remains “something went wrong”; with a log, there's the place to go to.

What to record about every solved error:

Entry partExampleWhat for
timeOctober 13, 2:32 PMdo errors pile up at a particular time (peak hour)?
inputthe message's type and the order numberwhich input led to the error?
errora timeout after 20 secondswhich type of error was it?
path takenretry — the second attempt succeededwas the main path, a retry, or the fallback enough?

A simple rule for alerting a human: alert when the fallback fires, and when the same error repeats a set number of times (for example five times a day). One retry needs no human; a fallback and a repeating error mean something has changed — someone has to know. In more depth, 4.6 Monitoring in production covers.

In plain terms: the log is a driving diary: three timeouts on Wednesday, all solved with a retry, are three entries — not three crises. When the diary shows that timeouts pile up around midday, you have a fact to decide on, not a feeling.

Step-by-step example: three incidents in the e-store

2.5 built the HomeCraft Store's return-request flow and tracked the first week with three numbers. The same week's three incidents:

#IncidentWhat happenedWhich path the system tookResult for the customer
atimeoutthe model answered in 30 seconds — the request exceeded the 20-second wait timeretry: the system tried again and the second time the answer came within secondsthe customer's message wasn't lost: the draft reached Piret for approval a minute late — for the customer an ordinary confirmation chain, just a small delay
bbroken outputthe model answered in free text (“Sure! The message's type is a return…”) — the automatic rule found no fieldsfallback: the check caught the breakage, the system took the standard answer and put the message on Piret's listthe customer immediately got a correct answer, the real answer within the promised day
cexternal errorthe email service was down for 10 minutes — the reply messages didn't go outsaving the work: the records waited in the queue and went out when the service recoveredthe customer got the answer about 10 minutes later — knows nothing about it

Three errors — and none reached the customer as an error. The work left was three entries in the log and one decision: the timeout's wait time was extended, because the timeouts piled onto the same half-day — the decision was made from the diary during working hours, not at night.

In plain terms: the first week gave three errors and zero emergencies — planned-for errors stay entries in a log and sometimes one small decision, not a customer's problem or a team's night.

Summary

  • Five error types, each with a usual solution — nothing needs to be invented in the middle of a crisis.
  • Retry only for temporary errors and with a stopping time (for example three attempts, the gap lengthening): rapid-fire retrying adds load to overload and becomes an error itself.
  • The fallback is the designated B path — a standard answer, saving the work, or a turn to a human — and always better than a wrong answer to the customer.
  • The log turns an error into a lesson: time, input, error, and path taken; a human is alerted when the fallback fires and on a repeating error.
  • Errors are normal and planned for — not an emergency: a system with paths ready lives through every error as one entry.

What's next?

  • previous → 3.3 Tools and actions
  • next → 3.5 Safety: limits and human-in-the-loop
  • Monitoring in production → 4.6
  • back → handbook index

Last updated 2026-10-05

← PreviousTools and actions: let the model actNext →Safety: limits and human-in-the-loop

© 2026 Siim Liimand · SeoWeb

GitHub/AI Handbook/Tallinn, Estonia

59.4370° N, 24.7536° E — Tallinn, Estonia

↑ Top