SeoWeb
  • Work
  • Services
  • CV
  • Contact
    • AI
  1. Home/
  2. AI Handbook/
  3. Agents & Evaluation/
  4. Production monitoring

[ Agents & Evaluation ]

Production monitoring

Target audience: technical + deep-diving non-technical | Prerequisites: 4.5 Evaluation: how to know whether a system is good

What you'll learn

After this document you will be able to:

  • explain what monitoring is (watching the system in production) and why a system that worked well yesterday can work worse today;
  • read the six core metrics (the “metrics” of 4.5; we now watch some of them continuously in production) and say which number is a signal and which is noise;
  • set up an alert (an automatic notification when a metric goes outside its allowed range) with a specific alert threshold (the limit whose crossing triggers the notification);
  • run a 15-minute weekly routine that feeds the test set (4.5) and prompt updates (2.6).

In plain terms

A system isn't launched once and ready forever: internal data changes, customers write new types of emails, and the model provider updates models. Monitoring means regularly looking at a few numbers that notice such a change before customers do. To start, six numbers, one specific alert, and 15 minutes a week are enough.

Evaluation vs monitoring: the planned test and reality

Quality can fall over time even when you change nothing yourself. Three typical causes:

  1. Internal data changes. Products are added to the catalog, prices swap, fields get a new shape — the flow that passed with the old data errs with the new.
  2. Customers write new types of emails. A new product category brings emails in language the classifier wasn't taught.
  3. The provider changes models. Updates arrive in the background — the answer's speed, language, or form can shift without you clicking anything. 5.5 covers this thoroughly.

All three are silent: the system doesn't send an error message, it simply starts doing slightly worse work. Monitoring is the only way to notice before customers do.

The difference from evaluation fits in one sentence: evaluation (4.5) is a planned test on known inputs; monitoring is the system's behavior on real inputs on a real day, when nobody is watching.

In plain terms: evaluation is the annual checkup at an agreed, scheduled moment; monitoring is taking your temperature every day. One tells how the system was at test time; the other, how it's doing in real life. Both are needed — one doesn't replace the other.

What to watch: the six core metrics

At the end of 2.5, the first week was counted with three numbers. Here the same idea scales into systematic monitoring — six metrics, enough for most systems:

MetricWhat it tellsWarning sign
volumehow many requests the system processed per daya sharp rise or fall = something changed: a new customer or a campaign — or the trigger is broken and the emails stop arriving at all
error rate (share of failures)how many % of requests ended in error — a timeout, corrupted output, an external failure; the entries come from the 3.4 loga rise above the usual level: something changed in the input, the form, or an external service
human escalation rate (how many % of cases went to a human)how many % of cases went to a human and how many % of approvals were changed before sending (see the 4.5 metrics)a rise on either count = quality has dropped or the input has changed
latencyhow long from input to answerslowing points to a provider change or a bottleneck — the technical side in 4.7
costhow much the system spends per day (3.6)a jump without a volume increase = something is running uncontrolled
special caseshow many emails went to the fallback path or to a humana sharp rise: the 3.4 rule — alert when the fallback fired and the error repeats

The first three numbers come from the same entries the flow already wrote to the history in 2.5 — monitoring isn't building a new system, it's regularly reading the existing entries. The metrics can be gathered onto a dashboard (all the numbers on one page at one glance), but as a first step a weekly summary is enough — in most no-code tools it exists as a ready-made template.

In plain terms: the six metrics are the system's vital signs, like pulse and temperature at a visit. Volume tells whether the patient showed up at all; the error rate, whether it was understood; the escalation rate, whether the work went well. Each alone is half the story — together they give the picture.

Alerts: specific, not for everything

If someone has to look at the numbers manually every day, the very day something happens is the day it gets missed. That's why the heart of monitoring is the alert: the system notifies by itself when an indicator crosses its allowed limit.

An alert lives and dies by its threshold:

  • A good alert is specific: “error rate above 10% per hour”, “the fallback path fired three times in a day”, “escalation rate above 20% per week”. Each such notification means a real change and deserves a look.
  • A bad alert is on everything: a notification for every error and every “missing” field. 40 emails a day means some small error every day — within the first week comes alert fatigue: so many notifications arrive that the human stops reading them, and the real alert goes unnoticed.

Who gets the notification? The maintainer — the role entrusted with the system's staying alive (1.6 defines it); one specific person, not a shared inbox where the message drowns. How an alert saves the day in real life is shown by the online store's Tuesday below.

In plain terms: an alert must be a smoke detector, not a doorbell. The doorbell rings for every guest and soon nobody rushes to the door; the smoke detector rings only for fire — and then people rush. Ten precise alerts a year are worth more than a hundred nuisances a month.

The weekly routine: 15 minutes that save

Alerts guard the nights and days; the human gets a weekly routine — 15 minutes in which the 3.4 log entries are looked at with three questions:

  1. What kinds of errors were there this week? Timeouts, corrupted outputs, external failures? The repetition of one kind points to one specific place to fix.
  2. Which entries went to a human and why? Did the classifier err, was data not found, or was the email genuinely unclear? A repeated error is a fault in the instructions, not bad luck.
  3. What's new in the customers' questions? New products, new words, new requests — everything the system doesn't yet know.

The answers go straight to work: every caught bug goes into the test set as a new case (4.5) and every new email type gets examples added to the prompt (2.6). This way the system grows from real input, not guesses. When there are several flows and the team is bigger, the same routine grows into a continuous improvement cycle — 5.6 looks at that.

In plain terms: 15 minutes a week is gardening: one walk along the fence to see what overgrew, what withered, and what's new. There is only one condition — the walk happens every week, not when something has already gone wrong.

A step-by-step example: the online store's first month

The home goods store's return flow (2.5) has been running for a month; after the 4.5 evaluation the human escalation rate was 12% and the error rate ~2%. What follows is visible only thanks to logging (the recording of events, 3.4): every email leaves an entry. Three things happened in the first month.

(a) Tuesday: the error rate jumped 2% → 14%. All the errors were timeouts (the 3.4 case type): a request exceeded the 20-second wait. The alert “error rate > 10% per hour” reached the maintainer before noon; the log's timing revealed the cause — the provider had updated response times overnight. The number of retries was increased and within a week the error rate was back at 2%. The alert saved the situation: without it the first sign would have been Piret asking “why are the approvals running late?” — and the one to notice last would have been the customer.

(b) The human escalation rate rose 12% → 19% over three weeks. Looking at a single week, the jump wasn't visible and the alert didn't fire — the rise was quiet. The weekly routine noticed it: the classifier didn't recognize a new product category (candle holders) it hadn't been taught, and those emails went to Piret as “unclear email”. Candle-holder cases were added to the test set and examples were added to the classifier's instructions (the 4.5 golden set; the 2.6 prompt update). The next week the rate fell to 13%.

(c) Cost rose 30%. First reflex: something is running uncontrolled. The dashboard comparison showed volume had grown by exactly the same amount — new customers bring more emails. Cost grew with volume; quality stayed the same. No action needed: this isn't a fault, it's the system growing.

MetricBeforeAfterAction
error rate2%14% (on Tuesday)alert → cause: the provider changed response times → retries increased; back to 2% within a week
human escalation rate12%19% (over three weeks)the weekly routine caught the cause: a new product category → test cases and examples added; 13% the next week
costbaseline+30%compared with the volume growth: new customers — normal, nothing done

Three events, three different paths: one through the alert, one through the routine, one turned out not to be a fault. None grew into a customer problem, because each was noticed in the numbers — not in a customer's email after the fact.

In plain terms: six numbers and one well-set alert turned the first month's three surprises into three calm decisions. The same month without monitoring would have produced a Piret wondering why the approvals are stuck, and a few customers whose email went unanswered.

Summary

  • Quality falls quietly: data changes, customers write new things, the provider updates models (5.5) — monitoring is the only way to notice before customers do.
  • Evaluation is a planned test, monitoring is reality — both are needed; one doesn't replace the other.
  • Six metrics are enough: volume, error rate, human escalation rate, latency, cost, special cases — the first three come from the same log entries the flow already writes.
  • An alert must be specific: a threshold with a number, not a notification for everything — otherwise alert fatigue sets in; notifications go to the maintainer.
  • 15 minutes a week is enough to start: three questions over the log, the results into the test set and the prompts.

What's next?

  • previous → 4.5 Evaluation
  • next → 4.7 Performance and latency
  • Continuous improvement in large systems → 5.6
  • back → handbook index

Last updated 2026-10-05

← PreviousEvaluation: how to know whether a system is goodNext →Performance and latency

© 2026 Siim Liimand · SeoWeb

GitHub/AI Handbook/Tallinn, Estonia

59.4370° N, 24.7536° E — Tallinn, Estonia

↑ Top