SeoWeb
  • Work
  • Services
  • CV
  • Contact
    • AI
  1. Home/
  2. AI Handbook/
  3. Agents & Evaluation/
  4. Evaluation: how to know whether a system is good

[ Agents & Evaluation ]

Evaluation: how to know whether a system is good

Target audience: technical + deep-diving non-technical | Prerequisites: 4.4 Multi-agent architectures

What you'll learn

After this document you will be able to:

  • explain why “looks good” is not a basis for decisions and what evaluation is (the planned measurement of a system's quality in numbers);
  • build a test set (a prepared collection of inputs that gets run after every change) and a golden set (human-approved correct answers);
  • read the six core metrics — from accuracy (how many % of answers are correct) to the human escalation rate (how many % of cases went to a human);
  • catch a regression (something that used to work now failing) before a change reaches customers;
  • say who evaluates and when: automation, the human sample review, and three fixed moments.

In plain terms

“Looks good” is an opinion, not proof. Instead, collect 30–50 real cases, write the correct answer next to each one, and run them after every significant change. The old version got 37 right, the new one 38 — the change is good. The new one got 35 — a regression you see on paper, not in a customer's email.

“Looks good” is not a metric

The system is live, the demos ran nicely, and the team says: “it looks pretty good”. Later it turns out that three people meant three different things by “good” — one meant beautiful answers, another speed, a third their own successful cases.

Three reasons why a feeling doesn't work as a metric:

  1. Everyone looks at different cases — the satisfied employee saw exactly those emails that went well.
  2. The model is variable — the same input doesn't always give exactly the same output, so one successful attempt means little.
  3. Nothing warns of decline — the provider updates models in the background and real inputs change; quality can fall so quietly that the customer notices first.

The solution is what 2.1 already taught for a single prompt: acceptance criteria (an agreed requirement whose fulfillment lets the result pass) and tests — here they are scaled up to evaluating the whole system: numbers before and after every change.

Measurement is not a technical detail, but a way to decide without arguing.

The golden set and the test set

Two concepts you will use constantly from here on:

  • The test set is the exam: a fixed set of cases that go through always the same way, after every change.
  • The golden set is the answer key: for every case it records which answer or decision is correct — and a human confirms it, not the model.

Evaluation essentially means: run the test set and compare every answer against the golden set. Building goes like this — one thorough day, not a separate project:

  1. Collect 30–50 real cases. Old cases (data anonymized) are powerful — they come from reality, not imagination. Pick them from different types: typical; borderline, where even an expert hesitates; poor in data and exceptional; and one where someone has tried to mislead the system.
  2. Write the correct answer or decision next to every case — for a classifier the correct type, for a summary the list of facts the answer must contain.
  3. Have an expert confirm them. The golden set is golden only if someone who knows the subject has reviewed the “correct” answers.
  4. Keep the set alive. Every bug caught in production goes into the test set as a new case — the set grows with the knowledge.

In plain terms: the test set is the exam, the golden set is the teacher-approved answer key. Quality is not “I like it” — it is a number: how many questions out of forty were answered according to the answer key.

What to measure: the core metrics

Six metrics (a metric is a measurement expressed as a number) give a sufficient picture in most systems:

MetricWhat it measuresExample
Accuracyhow many % of answers are correct38/40 types correct = 95%
Classification errorswhich type is mistaken for which2× “complaint” → “unclear email”
Format correctnesswhether the answer is in the agreed, machine-readable formJSON valid, fields present (see 2.2)
Latencyhow long from input to answeraverage 4 s, p95 12 s (see 4.7)
Costhow much processing one case costs€0.004 / email (see 3.6)
Human escalation ratehow many % of cases went to a human12%, target below 15%

Two explanations:

  • Classification errors. The simple idea of a confusion matrix (a table of errors with correct and predicted types crossed): rows are the correct types, columns the predicted types, and the diagonal holds the correct answers — without it you fix blindly, because accuracy can stay the same while the errors have moved to a more important type.
  • Human escalation rate. A too-high number means automation isn't working — the human is doing the work the system was supposed to do. A too-low number isn't always a win either: perhaps too much judgment has been baked into the system and risky cases go unreviewed. The right number depends on the price of an error (see 1.2).

Who evaluates? Three layers:

  1. Automated checks — format and rules: is the JSON valid, are the fields present, are unordered extras absent. Cheap and present on every run.
  2. The human sample — for example reading 20 answers a week; this is the reviewer's (see 1.6) routine work: review a sample, mark the errors, and take them into the test set. If the system stays in human-in-the-loop mode (a human approves the result before it is used), the sample can be taken directly from the review work.
  3. Model as judge — a bigger model can grade a smaller model's answers (“were the rules followed?”), but the judging model errs too: its verdict is never the final decision.

Evaluation is a planned test on known inputs; what happens continuously in production is watched by monitoring (4.6) — both are needed, but they are not the same thing.

Regression: evaluating changes

A regression is when a case that passed correctly before fails after a change. It is the nastiest risk, because a change is always made with a good intention: someone fixes case X and doesn't notice that case Y broke along the way.

Hence one firm rule: after every change — a new prompt, new instructions, a new model, or a new model version — run the test set before rollout and put the result in a table:

ChangeBeforeAfterDecision
two examples added to the instructions86%94%adopted
prompt reworded “for the better”94%90%rejected — a regression
new model version90%90%not adopted: improves nothing but is more expensive (3.6)

Three moments when to evaluate:

  1. Before launch — a full run: the system doesn't go live until the whole test set passes the acceptance criteria (2.5 showed how; 2.5's seven-case trial run checked the flow's launch; here the set is bigger, because we measure accuracy numerically).
  2. After every significant change — a regression test: we measure before, measure again after, and decide by the numbers.
  3. Periodically — a sample of real cases: for example once a month review a random selection — reality changes even when your system doesn't.

In plain terms: every change always improves something and may accidentally break something else. The test set is a photo before and a photo after — if something that used to be good is broken in the new photo, you see it in the table, not in a customer's email.

The complete procedure is a topic of its own — see 5.5; the core is simple: before the swap run the test set, after, run it again.

A step-by-step example: a new model for the old classifier

The home goods store's return flow (2.5) has grown: instead of three types the classifier now distinguishes 8 types — “return”, “info”, “complaint”, “legal claim”, “order change”, “address change”, “invoice question”, and “unclear email”; complaints and legal claims go straight to Piret with a priority flag. The model provider announces a new version — we run the evaluation.

Step 1 — the golden set. 40 real emails (data anonymized): 5 from each of the 8 types, each email with the correct type recorded and confirmed by the reviewer.

Step 2 — the old model's baseline. Run over the 40 emails and compared with the golden set: 37/40 correct (92.5%). The errors:

ErrorHow manyConsequence
“complaint” → “unclear email”2the complaint still reaches Piret, but without the priority flag
“return” → “info”1the customer gets an info draft instead of a return draft

Step 3 — the new model and regression. New version: 38/40 (95%) — but the error table shows what the summary number hides:

ErrorHow manyNote
“complaint” → “info”1REGRESSION — this email was correct on the old model
“return” → “info”1the same error remained
ChangeBeforeAfterDecision
new model version37/4038/40adopted; the “complaint” type's instructions are refined

One email that previously reached the complaints list correctly now got an info draft — the regression is real. Decision: the new model stays (95% vs 92.5%), the complaint instructions were sharpened with clearer examples, and the test set was run again: 39/40. The remaining “return → info” error was recorded as a case to watch.

Step 4 — the remaining metrics. Format check: 40/40 valid JSON with all fields — 100%. Human escalation rate in the first week of real use: 12%, the target was below 15% — the system does the work and Piret manually reviews about every eighth email.

Summary

  • “Looks good” is not a metric — quality must be expressed as a number and measurement must be repeatable: the same method, the same cases, numbers before and after.
  • The test set is a collection of inputs, the golden set is human-approved correct answers — 30–50 real cases from different types are a good start; every caught bug goes into the set.
  • Six core metrics: accuracy, error type, format, latency, cost, and the human escalation rate — both extremes of the last one are warnings.
  • After every change run the test set — regression is caught in a table (change | before | after | decision), not in customer emails.
  • Who and when evaluates: automation on every run, the human sample periodically, a model used as judge only as an aid; evaluation happens before launch, after every significant change, and once a month.

What's next?

  • previous → 4.4 Multi-agent architectures
  • next → 4.6 Production monitoring
  • back → handbook index

Last updated 2026-10-05

← PreviousMulti-agent architecturesNext →Production monitoring

© 2026 Siim Liimand · SeoWeb

GitHub/AI Handbook/Tallinn, Estonia

59.4370° N, 24.7536° E — Tallinn, Estonia

↑ Top