SeoWeb
  • Work
  • Services
  • CV
  • Contact
    • AI
  1. Home/
  2. AI Handbook/
  3. Enterprise Scale/
  4. Continuous improvement: from measurement to decisions

[ Enterprise Scale ]

Continuous improvement: from measurement to decisions

Target audience: non-technical leadership + technical | Prerequisites: 5.5 Model updates and drift

What you'll learn

After this document, you'll be able to:

  • explain the principle “measurement itself doesn't improve anything — the decision does” and describe the improvement cycle (data → analysis → change → measurement);
  • run a monthly improvement cycle in five stages — following the table, with each stage's output in view;
  • prioritize ideas with the impact × effort grid and boldly write down “don't do” as well;
  • find improvement ideas from five sources, the most valuable of which is the list of human interventions;
  • also make the decision to end (5.4) and keep a small, regular rhythm for the cycle.

In plain terms

Monitoring (4.6) gives you numbers, but a number doesn't improve anything — the thing that improves is a person who, after looking at the numbers, decides to do something differently. Continuous improvement is an agreement: once a month the team sits down with the numbers, picks one change, writes down the expected effect, and checks a month later whether it happened. Two hours a month gives you twelve measured improvements a year — more than one big annual review.

Measurement doesn't improve — the decision does

4.6 ended with a weekly routine: 15 minutes a week over the log with three questions. That routine collects data and catches alerts, but so far nothing has been improved. A metric (the measurable indicator introduced in 4.5 and 4.6) is only a message; the improvement is born from a decision: what we change, what we expect from it, and how we know that it happened.

If the team looks at the numbers month after month and nothing changes, monitoring is no longer a tool — it's a ritual. That's why every measurement ends in a decision, and the decision leads to a new measurement — this is called the improvement cycle:

   ┌─────────────────────────────────────────────┐
   ▼                                             │
 data ──► analysis ──► change ──► measurement ───┘
  • data — the monthly report; analysis — what repeated; change — one at a time, tested; measurement — did the metric move?

The cycle runs as long as the system lives — but only if every loop ends in a decision. Without a decision, the cycle remains a drawing.

In plain terms: the scale doesn't make you lose weight — the scale only shows whether what you're doing works; measurement says “improved or not,” but it's the decision to change something that sets things in motion.

The monthly improvement cycle in five stages

The weekly routine (4.6) stays — its job is collecting data and catching alerts quickly; for the bigger work the same idea grows into a monthly cycle: once a month, 1–2 hours, a fixed day and participants. Dividing the roles is the team's agreement — taught by 5.2 Team workflows and standards.

StageWhat is doneWhoOutput
1. Collecting the datathe monthly report: 4.6's six metrics, 3.4 log entries, the list of human intervention casesmaintainerthe month in numbers + the list of interventions
2. Analysistwo questions: what consumes the most human time? which errors repeat?content designer + reviewera list of recurring problems with frequencies
3. The decision listevery idea written down: what we change and what's expected — the expected effect on the metricthe whole team, the sponsor decidesthe decision list — ideas with expected effects, prioritized
4. The improvementone change at a time (2.1), a regression test before going to production (4.5)builderan updated prompt or knowledge base + the test result
5. The check a month laterdid the metric move in the expected direction?maintainer + sponsorthe decision: kept / reverted / the next idea

Two places where the cycle most often breaks:

  • At stage 3, the effect isn't written down. “We'll improve quality” isn't an expected effect; “intervention rate 14% → below 10%” is. Without a number, it's impossible a month later to say whether the change succeeded.
  • At stage 4, several changes are made at once. 2.1's principle applies here too: one change at a time, otherwise you won't know a month later which one helped. Regression (the failure of something that used to work, the 4.5 principle) is caught by the test set before customers.

If you're the manager who leads the monthly review, your job isn't reading the tables aloud but asking questions:

  1. Which metric moved the most this month — and why that one?
  2. How many hours did people spend on interventions? Did the same cases repeat?
  3. Last month's change: did the metric move as promised in the decision list? If not — why?
  4. What's new among the customers that the system doesn't yet know?
  5. Which ideas went unmade this month — and are they still important?

Where the ideas come from: five sources

  1. Human interventions — the most valuable source. Every manually corrected answer is a lesson: if the reviewer changed the answer before sending, they know exactly what was wrong. The list of interventions gives concrete places to improve — free expert judgment you can't buy elsewhere.
  2. Customer feedback. Recurring questions and pain points straight from the customer: what stayed unclear, which answer helped, which didn't.
  3. Alerts (4.6). Every fired alert is a half-analyzed case — all that's left is to find the cause and decide whether to fix it.
  4. The prompt bank's history (2.6). The version history shows what has already been tried and what happened then — many “new” ideas are actually old ideas nobody remembers.
  5. Test cases that failed (4.5). Every failed case describes exactly the situation the system can't yet handle — and stays afterward as a check that the regression doesn't return.

The decision to end

Continuous improvement also includes the courage to end: when the benefit/cost comparison (5.4) shows for several months in a row that maintaining the system costs more than it saves, the right decision is to take the system down — not to keep improving a thing nobody needs anymore.

The rhythm: small and regular brings down the big one

The monthly cycle is dental care: a small regular step, not an annual root canal. A big review done once a year stacks up twelve months of changes in advance: the data has changed several times in between and the numbers can no longer be separated — nobody knows what had an effect and what didn't. Two hours a month keeps every change small and reversible.

Prioritization: impact × effort

Impact × effort is prioritization's simplest grid. For every row of the decision list, two questions: how much does it move the metric? and how many work hours does it cost?

Low effortHigh effort
High impactdo it now — this monthplan — take it into next month's cycle and split the bigger work into parts
Low impactdo it on the side if time remains — or skipdon't do

The most important square is “don't do.” Every cycle produces ideas that seem sensible but move no metric. Writing that square down is as important as doing — a month later, nobody has to ask why it wasn't done.

In plain terms: ask two things of every idea — “does the metric move?” and “how many work hours does it cost?” — and half the list falls away by itself. The four squares keep a month's work from being spent on the wrong thing.

Step-by-step example: CallHelp's three months

The CallHelp chatbot — the same phone support where 2.4 cleaned transcripts and 4.3 built the customer profile — has been in production for three months (4.7 meanwhile made the bot three times faster). The main metric is the human intervention rate — what % of chats goes to a customer service agent (the 4.6 principle).

Month 1 — data and the decision list. The intervention rate is 14%. Three recurring errors rise from the list of interventions:

  1. modem models — the bot mixes up the devices' models with one another and gives wrong instructions;
  2. dates — the customer asks “will it be solved by Wednesday?”, but the bot doesn't know what date Wednesday is and promises something else;
  3. the greeting format — answers start one way one time and another way the next; some are awkwardly long.

Three recurring errors, three rows in the decision list:

ChangeExpected effectPriority
label refinement: clear rules for models and dates in the prompt + 5 new test casesintervention rate 14% → below 10%high impact / low effort → now
customer profile field update: the profile receives the current date and installation datareducing the same error, but it needs development workmedium impact / medium effort → plan
shortening the answer templatethe greeting is form, not content — it doesn't change what the customer thinks of the resolutionlow impact → don't do

Month 2 — the change and the check. The builder did only the first row — one change at a time (2.1) — and ran the test set before going to production (4.5): no regression appeared. By the end of the month, the intervention rate was 9% — the expected effect was met.

Month 3 — a new phenomenon. The weekly routine (4.6) noticed something new: customers now ask about fiber internet installation — a topic missing from the knowledge base, so those chats went to a human. An installation chunk was added to the knowledge base (4.2) along with five test cases as a check. Whether the new topic stays will be shown by next month's stage 5.

Just as important is what CallHelp didn't do: the template shortening stayed undone — and nobody asked for it; the profile update lives in the decision list and will come when planned. Three months, one measured improvement, one planned piece of work, and one written “no.”

Summary

  • Measurement doesn't improve — the decision does: a metric is a message; the change is born from the decision “what we change and what we expect.”
  • The cycle has five stages: data → analysis → the decision list → one change (with a regression test) → the check a month later.
  • Every decision-list row must carry a numeric expected effect — otherwise you can't say a month later whether it succeeded.
  • Impact × effort prioritizes: high impact and low effort now, high impact and high effort plan, low impact and high effort — don't do.
  • The best source is the list of human interventions: every manually corrected answer is a concrete lesson you can't buy elsewhere.
  • Small and regular brings down the big one: two hours a month keeps the changes small and reversible — and includes the courage to end (5.4).

What's next?

  • previous → 5.5 Model updates and drift
  • next → 5.7 Responsibility, ethics, and governance
  • back → handbook index

Last updated 2026-10-05

← PreviousModel updates and driftNext →Responsibility, ethics, and governance

© 2026 Siim Liimand · SeoWeb

GitHub/AI Handbook/Tallinn, Estonia

59.4370° N, 24.7536° E — Tallinn, Estonia

↑ Top