- Home
- AI Handbook
- Enterprise Scale
- Template library: checklists and samples
[ Enterprise Scale ]
Template library: checklists and samples
Target audience: everyone | Prerequisites: the whole handbook (no single document is a required prerequisite — this is a library to use as needed)
What this library is and how to use it
This is the handbook's last document — and essentially a toolbox. The fifth level taught how large projects can be organized; this library gathers that for copy-and-paste use: four checklists (a checklist — a fixed list of points walked through before moving on) and four templates that refer back to what the handbook taught. There's no new theory here — only a summary in usable form.
Using it is three steps:
- Take a template. Choose by the situation: a new system going live → checklist 1; a prompt change → checklist 2; the month has ended → checklist 3; a model change → checklist 4; documentation → the templates below.
- Fill it in. Copy it into your own document and go through the points. The mark “☐” means “not done” — next to a done point, add the date and a name.
- Write it down. A filled-in checklist isn't one-time paper: it stays with the system's documentation as evidence of what was checked, when, and by whom.
In plain terms: this library is like a pre-flight checklist: a pilot doesn't read it because they couldn't manage without it, but so that no step slips their mind — and so the log afterward shows who checked and when.
Checklist 1: a new system before going live
The birth protocol (the check a new system passes before going live, 5.2): before the first result reaches a customer, every point must be ticked.
- ☐ The task has passed the decision mask — you know why exactly this task suits an AI and where a human remains irreplaceable → 1.2 Capabilities and limits
- ☐ The safety test done in writing: what is the worst thing the system can do, and which safeguard (a limit built into the design, not hope in a well-behaved model) prevents it → 3.5 Safety
- ☐ Human-in-the-loop (a workflow where a human approves the result before use): approval and stop points written down → 3.5 Safety
- ☐ Golden set (human-confirmed correct answers) ready, approved by an expert, the system's results written into a table → 4.5 Evaluation
- ☐ Cost estimate made and the alert threshold (the cost limit whose crossing triggers a notification) set → 3.6 Cost management
- ☐ Security reviewed: API keys protected, data access restricted, prompt injection blocked → 3.7 Security
- ☐ Architecture sheet: the system's parts and dependencies on one page → 5.1 Architecture at scale
- ☐ Owner (the person responsible for a system/prompt, a single name) designated and written down → 5.2 Team workflows and standards
- ☐ Ready for monitoring: metrics and concrete alerts up, notifications go to one person → 4.6 Monitoring
- ☐ A prompt bank entry exists (template below) → 2.6 Prompt management
The rule is simple: if any point stays empty, the system doesn't go live — the first result that reaches a customer must not also be the first test.
Checklist 2: a prompt change in production
A prompt in production isn't a file to be corrected on the fly — a change is a piece of work that always goes through the same flows (2.6, 2.1, 5.2).
- ☐ The need for the change written down: what is being changed and why (into the log, not into memory)
- ☐ Test cases ready: a sufficient selection from the test set, including edge cases → 4.5 Evaluation
- ☐ Comparison against the golden set: numbers before and after into a table — “seems better” isn't a result → 4.5 Evaluation
- ☐ A second review: one more person has looked over the change → 5.2 Team workflows and standards
- ☐ Approval: only the prompt's owner or a named deputy sends it to production → 5.2 Team workflows and standards
- ☐ Bank: a new version number, the reason for the change recorded, the old version to the archive → 2.6 Prompt management
Checklist 3: the monthly review
Once a month — 30 minutes to fill in the table; the whole cycle with discussion 1–2 hours, where the monitoring numbers become decisions (the routine from 4.6, the cycle from 5.6). Fill in the table:
| Metric | Last month | This month | Note |
|---|---|---|---|
| volume (requests per day) | |||
| error rate | |||
| human intervention rate | |||
| latency | |||
| cost | |||
| special cases (routed to fallback/to a human) |
Explanations of the metrics and the warning signs are in 4.6 Monitoring. Then three questions:
- Which metrics moved and why? A sudden change without a cause is worth investigating — even when it's in a good direction.
- What went to a human and why? A recurring error is an instruction bug, not bad luck → 4.6 Monitoring.
- What do I do with this during the month? Every caught bug goes into the test set (4.5), every new email type goes next to the prompt as an example (2.6) — this is how measurement is about decisions, not merely collecting numbers (5.6).
Checklist 4: a model change
The six steps in brief (5.5 Model updates — applies to a planned migration and to an unplanned change alike):
- ☐ Dependencies marked down: who uses the old model and whom the change affects → 5.1 Architecture at scale
- ☐ Golden set run against the new model, results into a table → 4.5 Evaluation
- ☐ Results compared: accuracy, latency, cost — before and after → 4.5 Evaluation
- ☐ Prompts refined: if the new model is weaker on some class, examples added → 2.1 Writing effective prompts
- ☐ Through the test environment, then to production → 5.1 Architecture at scale
- ☐ Followed up in the first week: numbers in real use, not only in tests → 4.6 Monitoring
In plain terms: a change is like moving house — you count your things, try out the furniture, compare the bills, and spend a week checking whether everyone manages in the new place.
Template: a prompt bank entry
Next to every prompt, the bank keeps six core data points + a testing reference (2.6's template has six parts; the library adds the testing reference); fill in:
PROMPT BANK ENTRY
-----------------
Name: customer-email-classifier
Purpose: Sorts customer emails into 8 classes (return, info, complaint, …) (see 2.5 and 4.5)
User: the webshop customer-support workflow (step 1), daily
Last tested: 2026-10-01 — golden set 40/40
Status: in use (other values: testing | archived)
Version: v3 (2026-09-28 — examples added for the "complaint" class)
Testing reference: the 4.5 test set, the table "classifier v3 results"
The “last tested” date is the most important line: it tells you the age of the knowledge.
Template: the decisions table and risk register
When standards grow to company-wide scope, it becomes a governance (the system's steering rules) question (5.7). Two templates to start with:
DECISIONS TABLE
---------------
| Decision | Who makes it | Who approves | Where it's written |
|------------------------------|---------------------|-------------------|----------------------|
| New system go-live | sponsor + builder | manager | birth protocol page |
| Prompt change in production | builder | prompt's owner | prompt bank log |
| Model change | maintainer | owner | decision log |
| Alert threshold change | maintainer | system owner | architecture sheet |
RISK REGISTER
-------------
| Risk | Probability | Impact | Who watches |
|-------------------------------------------|-------------|----------|-------------|
| Provider changes the model without notice | medium | high | maintainer |
| Costs grow faster than the volume | medium | medium | owner |
| Prompt injection through user input | low | high | owner |
The rule for both: one name next to every row, not a committee — otherwise responsibility defaults to nobody (5.7).
Template: the decision protocol
Every significant decision has a written protocol (5.7): what, why, who, when.
DECISION PROTOCOL
-----------------
What was decided: [the decision]
Why: [the reasoning, based on data]
Who decided: [name, role]
When: [date]
Closing word: the handbook is complete — what's next?
With this, “The AI Automation Handbook” — 34 documents across five levels — is complete. The best next step isn't more reading but picking one service and building it end to end: take one tedious, repetitive task, run it through the decision mask (1.2), and build the first workflow from start to finish (2.5). Grow the system only after the first one works and has been measured — the levels must stay in order, because each next one assumes the previous. If something goes wrong along the way, this library is there for exactly that: checklist up, template filled in, result recorded. The whole handbook is reachable via the handbook index, and the first step is always 1.1 What is an AI model and how it “thinks”.
- previous → 5.7 Responsibility, ethics, and governance
- start from the top → 1.1 What is an AI model and how it “thinks”
- back → handbook index
Last updated 2026-10-05