- Home
- AI Handbook
- Enterprise Scale
- Model updates and drift
[ Enterprise Scale ]
Model updates and drift
Target audience: technical + non-technical leadership | Prerequisites: 5.4 Cost strategy at scale
What you'll learn
After this document, you'll be able to:
- distinguish two kinds of change: a planned model migration (adopting a new model) and an unplanned change (the provider updates the model underneath you);
- explain what drift (model drift — the model's behavior shifting over time without notification: same input, slightly different answer) is and why, in a large system, one model touch affects several places at once;
- walk a change through six steps — from the dependency sheet to follow-up monitoring;
- protect the system against an unplanned change: a pinned model version (a pinned exact form, e.g. v3), monitoring alerts, and a monthly test run;
- tell a manager why a change is a process backed by numbers, not a button press.
In plain terms
Your automation runs on another company's model, and that model doesn't stand still. Sometimes the provider announces in advance: “a new version is coming, the old one is leaving” — then you pick the moment and walk the change through with tests. Other times the provider quietly changes something — and the only one who notices is your monitoring. In both cases the rule is the same: know who uses which model, run the tests before the change, and watch the numbers after. A model change is a six-step process.
Two kinds of change: the planned migration and the unplanned change
A planned migration is your decision. You adopt a new model because it's cheaper (5.4), faster, or more capable. You choose the date, you run the tests, and the decisions come from the numbers (4.5).
An unplanned change is the provider's decision: it updates the model underneath you. 4.6's warning sign “the provider changes models” refers to exactly this — the response's speed, language, or form can shift without you clicking anything. The webshop saw it in the first month: the provider changed response times overnight and the error rate jumped from 2% to 14%.
The third, quietest kind is drift: a single case isn't a bug, but the shifts add up over weeks and months, and quality falls so calmly that the customer notices first.
Why this is serious in a large system:
- The dependency (which system uses which model) is a long list in a large system — the same model serves dozens of flows, agents, and several teams' work. One model change affects all users at once — and if the architecture sheet (5.1) doesn't keep that list, you won't know whom to notify and whom to put to testing.
- Prompts are tested against a specific model. The prompt bank keeps “last tested” and the version next to every prompt — but that test was done with a particular model version. Change the model and the whole bank is untested: examples and rules that worked with one model don't carry over to another.
| Kind of change | Whose decision | How fast | How you notice |
|---|---|---|---|
| Planned migration | yours | at the moment you choose | test numbers before and after |
| Unplanned update | the provider's | overnight | a monitoring alert |
| Drift | nobody's, quietly | over weeks, months | a quiet rise in the intervention rate |
In plain terms: a planned migration is your own house move — you pack yourself and pick the day. An unplanned change is the neighbor renovating your wall — no notice comes; the only remedy is measuring. Drift is a house sinking a centimeter a year: no single day is critical, but ten years are.
The change process in six steps
The process applies to both kinds: with a planned migration you pick the time; with an unplanned update the provider sets the deadline.
| # | Step | Action | Tool |
|---|---|---|---|
| 1 | Mark down dependencies | who uses the old model and whom the change affects | 5.1 architecture sheet |
| 2 | Run the golden set against the new model | the test set with the new version, the result into a table | 4.5 regression |
| 3 | Compare the results | accuracy, latency, cost — before and after | 4.5's six metrics |
| 4 | Refine prompts | if the new model is weaker on some class, add examples | 2.1 templates |
| 5 | Through the test environment, then to production | a pass through the isolated environment before real customers | 5.1 environments |
| 6 | Follow up during the first week | numbers in real use, not only in tests | 4.6 monitoring |
Two steps deserve their own sentences.
Steps 2–3 are the core. The golden set (human-confirmed correct answers, the 4.5 principle) gives a comparison that isn't opinion. Regression (the failure of something that used to work) often hides in the error table, not the summary number: one class improved, another broke. The decision is written like the table in 4.5: the change, before, after, and the decision.
Step 6 is the safety net. The test shows how the system behaved with 40 known emails; production shows how it behaves with thousands of real ones. With the intervention rate steady and no alerts, the change is done — with an alert, you know within a week, not from a customer's email.
In plain terms: a change is like moving house: you count who works where, try out the furniture, compare the bills, fix the plan, move first into the backup office, and spend the week checking whether everyone manages in the new place.
Protection against an unplanned change
A planned migration can be scheduled; an unplanned one can't. Three protections make the unplanned change manageable:
- Pin the model version when the provider allows it. The request doesn't say “give me the newest” but an exact identifier, e.g. “model-v3-2026-06”. The provider can't quietly swap anything — the version stays until you change it. Two limits: not every provider allows this, and support for the old version ends anyway. Pinning buys time; it isn't an eternal solution.
- Monitoring alerts (4.6) are your backup sensor. An alert threshold (the limit whose crossing triggers a notification) on a numeric indicator — e.g. “p95 latency above 10 s” or “intervention rate above 20% in a week” — catches the provider's change even when you didn't read the provider's newsletter.
- The periodic test run. Once a month, run the golden set through — even if nothing was changed. If the result jumps from 39/40 to 36/40 with no change of yours, the change happened underneath you. As a bonus, it keeps the prompt bank's “last tested” date (2.6) fresh.
The fourth protection is long-term — the exit strategy. Don't burn with a single provider: keep the abstraction (a provider-independent layer, the 5.1 gateway principle) — the model name and key live in one place in the system, replaceable with a single sentence. A documented golden set makes a provider change possible: the tests run through in either direction.
In plain terms: a pinned version is a deadbolt — an ordinary key no longer opens it. Alerts are the alarm that rings if someone still gets in. The test run is a monthly inventory — it shows whether the shelves are in order even when nobody has touched anything.
Step-by-step example: the webshop classifier v2 → v3
The webshop classifier's state after 4.5: eight classes, a golden set of 40 real emails, v2 in production with a result of 39/40, a human intervention rate of 12% (target below 15%), average latency 4 s and p95 12 s, cost 0.004 € per email. Then a letter arrives from the provider: v2 leaves in 60 days, v3 takes over — an unplanned change that forces a planned migration.
The six steps:
| # | Step | At the webshop |
|---|---|---|
| 1 | Dependencies | architecture sheet: the classifier uses v2 — one system, used by Piret and the customer support team |
| 2 | Golden set | 40 emails through v3: 38/40 — the error profile changed |
| 3 | Comparison | latency −40%, cost −15%, accuracy lower at first |
| 4 | Prompt refinement | two examples added to the “complaint” instructions → a new run gave 40/40 |
| 5 | Environments | a week in the test environment, then to production |
| 6 | Follow-up monitoring | a week in production: the intervention rate held at 12% |
The error table (step 2) shows what the 38/40 number hides:
| Error | v2 | v3 | Consequence |
|---|---|---|---|
| “return” → “info” | 1 | 0 | the case 4.5 had its eye on improved |
| “complaint” → “info” / “unclear email” | 0 | 2 | regression — complaints stay unflagged for importance |
The new model is better at “returns,” weaker at “complaints.” The regression is fixed in step 4: the complaint instructions got two clear examples (the 2.1 template), the new run gave 40/40, and the prompt entered the bank as a new version with a record of what was changed, why, and who approved. The decision, with numbers:
| Change | Accuracy | Latency | Cost | Decision |
|---|---|---|---|---|
| v2 → v3 + prompt refinement | 39/40 → 40/40 | average 4 s → 2.4 s; p95 12 s → ~7 s | 0.004 → 0.0034 €/email | adopted |
And then the unplanned change. Three months later an alert fired: p95 latency, which had settled in production to ~6 seconds (faster than the test week's ~7 s), jumped overnight to 9 seconds — without anyone changing anything. The pinned version number hadn't changed, the golden set still gave 40/40 — only the time had lengthened. The provider was asked — it turned out that part of the v3 traffic had been moved overnight to different infrastructure. The fix: the system was pinned to the exact version identifier “model-v3-2026-06”, and the provider promised to announce infrastructure changes in advance; the alert threshold stayed in place. Within a week, p95 was back around ~6 s.
In plain terms: the planned migration took two days of work and brought a 15% cost saving — because the tests already existed. The unplanned change cost one alert and one question to the provider — because the alarm rang. Both stayed small, because the protection had been built before, not after.
Summary
- Two kinds of change: a planned migration is your decision and your timing; an unplanned change is the provider's — 4.6's warning sign “the provider changes models”. Drift adds up quietly.
- The change touches the whole system: the dependency (which system uses which model) sits on the architecture sheet, and prompts are tested against a specific model version — a model change alters both at once.
- Six steps: dependencies → golden set through → comparison (accuracy, latency, cost) → prompt refinement → test environment and production → a week of monitoring.
- Three protections against an unplanned change: a pinned model version, alerts with alert thresholds, and a monthly test run.
- The exit strategy: the abstraction (a provider-independent layer, 5.1's gateway principle) and a documented golden set make a change possible — including a change of provider.
What's next?
- previous → 5.4 Cost strategy at scale
- next → 5.6 Continuous improvement: from measurement to decisions
- back → handbook index
Last updated 2026-10-05