SeoWeb
  • Work
  • Services
  • CV
  • Contact
    • AI
  1. Home/
  2. AI Handbook/
  3. System Architecture/
  4. Cost management: tokens, prices, budget

[ System Architecture ]

Cost management: tokens, prices, budget

Target audience: technical + non-technical readers who want to go deeper | Prerequisites: 3.1 API integrations

What you'll learn

After this document you will be able to:

  • explain where the price of AI usage comes from — input, output, and the growth of conversation history;
  • make a simple cost estimate (a cost forecast made before going live) before going live, with one formula;
  • find the real numbers on the provider's dashboard (the providers' web portal where actual usage and costs are visible);
  • name five practices that lower cost without sacrificing quality;
  • set an alert threshold (a cost limit beyond which a notice is sent) and a budget limit (a pre-set monthly cap).

In plain terms

Using an AI model is paid for for every question and every answer, piece by piece. Costs don't depend only on the price list, though — two things you control dictate them: how much you send along with each request and how often you ask. An estimate before going live, one look at the dashboard a month, and a sensible budget limit prevent big surprises.

How the price forms: tokens going in and coming out

From 1.1 What is an AI model and how it “thinks” you know what a token is (a piece of a word). In API use it's also the billing unit: on the request shown in 3.1 API integrations you pay in two places.

Input — everything that goes along with the request. Instructions, data, history — everything put into the context window (the amount of text the model sees at one time) is billed.

Output — the model's answer. In most price lists an output token is more expensive than an input one (typically a few times over) — writing the answer piece by piece is computationally more expensive than reading.

The third factor matters most to the system builder: conversation history (the collection of earlier turns that the system sends along again with every request). As 3.2 Context management explains, the model remembers nothing — on every request the whole history so far goes along again and is paid for again. So the fifth question of a ten-turn conversation is billed for ten turns' worth of input — an unchecked growing history grows the costs at the same pace.

In plain terms: you pay both for what you say and for what is said back. And since the model remembers nothing, it has to read out all the earlier talk every time — and pay for it again.

One hidden cost more: errors. When the system repeats a request, the retry is paid for too — a request repeated many times is many times more expensive; error-handling principles are 3.4 Errors and error handling's topic.

The cost estimate before going live

In plain terms: before going live, calculate the cost on paper — like a budget before construction. The result doesn't have to be cent-exact; it has to show the order of magnitude: is this talk of one euro a month or a hundred.

The cost estimate goes with one formula:

daily cost = requests × (input tokens × input price)
           + requests × (output tokens × output price)

A worked example:

  Requests per day:                12
  Average input:             2,500 tokens
  Average output:              400 tokens
  Sample price, input:    €0.002 / 1,000 tokens
  Sample price, output:   €0.006 / 1,000 tokens (more expensive!)

  Input:   12 × 2,500 = 30,000 tokens → 30 × €0.002 = €0.06
  Output:  12 ×   400 =  4,800 tokens =  4.8 × €0.006 = €0.029 ≈ €0.03
  Day total:                              ≈ €0.09
  Month (22 working days):                ≈ €2

The estimate here covers only the daily summaries; in the invoice-entry drafts flow we multiply the same formula by our own numbers.

Two honest remarks. First: the prices are illustrative and change. “€0.002 per thousand input tokens” is a teaching sample figure — you'll find the real prices in the providers' price lists; they depend on the model (big more expensive, small cheaper) and change over time. Second: round boldly — the goal is the order of magnitude, not the cent. The average token count per request emerges from one trial run or from the dashboard; the sponsor's decision also leans on that order of magnitude (see 1.6 Roles and responsibility in a project).

Monitoring: the dashboard, the review, the alert

In plain terms: monitoring (continuously tracking the system's operation and costs) doesn't have to be complex. Three habits: a glance at the dashboard every week, going over the numbers once a month, and one alert setting. More isn't needed at the start.

Where to see the cost: the dashboard. Every major provider offers a dashboard with the real usage's cost — by days and by models. For the technical reader: you can also see there how the tokens split between input and output and the costs of individual requests.

Once a month, over the numbers — the basis of the sponsor's decisions. By 1.6, the maintainer (in the Number Bureau, Jaak) tracks the cost pace; the sponsor Anu goes over the numbers against the estimate once a month.

The alert threshold. Dashboards let you install an alert threshold: for example “when the month's cost exceeds €7, the system sends Jaak an email”. The alert is an early sign: usage behavior has changed — before the bill surprises.

Five ways to reduce cost

In plain terms: all five ways are variants of three principles — send less along, ask fewer times, get a shorter answer. None requires programming; all require thinking.

  1. A smaller model for simpler tasks. Document 1.1's principle: the specialist is expensive, the intern cheap — classifying a message (“invoice” / “complaint” / “order”) doesn't need a big model; leave the big one for complex tasks.
  2. A shorter prompt (the instruction given to the model). The standing instruction spends tokens on every request — 2.1 Writing effective prompts' persistent parts of a prompt must be precise, not long. Don't repeatedly send a big instruction when a short one suffices.
  3. Fewer requests. The system shouldn't call the model on every keystroke — wait for the input's end, merge several small questions into one. Every call is a billable event.
  4. A shorter output. Ask only for the needed fields — 2.2 Structured output helps: instead of ten sentences, three numbers, and a shorter answer is cheaper.
  5. A shorter history. If the system carries the whole history along on every request, it pays for it again every time — the “state of affairs” solution (see 3.2) keeps the key facts in a separate short record and the window short.

One rule holding in all five: cheaper must not mean worse quality — after every saving step, check that the quality held. Deeper optimization — caching, batch processing, measuring latency and performance — 4.7 Performance and latency covers.

Budget and the budget limit

In plain terms: the budget limit is like a credit card limit — the whole month can be spent, but when the limit is full, spending stops and this is reported. A surprise bill stays away.

The estimate and the alert threshold give understanding; the budget limit adds protection. The monthly cap can be installed at three levels — and the choice is the sponsor's decision, because it's a budget question (see 1.6):

  • Notifies. The softest variant: when the threshold is exceeded, a notice is sent, the system keeps working.
  • Reduces. The system switches to a cheaper model or pauses non-critical work — for example summaries are made with the smaller one, client work stays unchanged.
  • Stops. The harsher variant: the budget limit full — new requests no longer go through. Suits only work where a pause is tolerable: a nightly report can wait, a customer's answer can't.

A sensible balance: a budget limit about twice the estimate (leaves room for growth), an alert threshold at about 70–80% of the budget limit. When usage has large volume and constant growth, cost-planning strategy is 5.4 Cost strategy at scale's topic.

Step-by-step example: the Number Bureau's monthly cost

Again the “Number Bureau” (see 1.6): owner Anu, maintainer Jaak. The system composes entry drafts from invoices and additionally every day gives 12 clients a daily summary of their invoice operations. Anu asks Jaak: “What will this AI cost us?”

1. The estimate on paper. Jaak does the above calculation for the daily summaries and multiplies the same formula with the drafts' flow — in total he arrives at ~€4 a month (of which ~€2 the daily summaries). “A couple of euros a month,” Jaak says. Anu is satisfied.

2. The first week on the dashboard. A week later Jaak shows Anu the provider dashboard's data for the past week:

Weekly cost:      ~€1.20   (monthly pace ~€5.20)
Input tokens:     ~490,000
Output tokens:    ~24,000

The amounts are small, but the ratio is wrong — why?

3. One prompt eats more than half. In the split a surprise appears to Anu: about 60% of the week's cost (~€0.70) comes from one single daily prompt — one client's summary. The cause is found quickly: in this workflow the system carries the whole week's conversation history along on every request — and as 3.2 teaches, the whole history is billed again on every request. Jaak applies 3.2's third strategy — the “state of affairs”: the system keeps the key facts (client, period, amounts) in a separate short record (~1,500 tokens) and the long history no longer goes along.

A week later on the dashboard: the week's cost ~€0.60, monthly pace ~€2.50–3 — the cost falls by half and is again in line with the estimate. Quality stays the same: the facts are now precisely recorded, not left to a long history's care (how to check that, 4.5 Evaluation teaches).

4. The protection in place. Finally Jaak installs an alert threshold of €4.50 a month (a notice to Jaak and Anu) and a budget limit of €6 a month (it is set higher than twice the estimate, because the real pace approached the estimate). If the cost were to exceed it, the system would first switch the daily summaries to the smaller model and notify — the entry drafts would stay unchanged.

The estimate gave the order of magnitude, the dashboard showed the actual, one review found one too-long prompt, and the protection keeps future surprises away. The whole exercise took less than an hour.

Summary

  • The price comes from tokens — in input and output alike, output is usually more expensive; conversation history is the main growth factor, because the whole history is billed again on every request.
  • Make the cost estimate before going live — a simple formula gives the order of magnitude; prices are illustrative and change, you'll find the real ones in the provider's price list.
  • Monitoring is three simple habits: a glance at the dashboard, the maintainer's monthly review, and an alert threshold.
  • Five things cut cost: a smaller model for simpler tasks, a shorter prompt, fewer requests, a shorter output, a shorter history.
  • The budget limit is protection with three variants: notify, reduce, or stop — the choice is the sponsor's budget decision.

What's next?

  • previous → 3.5 Safety: limits and human-in-the-loop
  • next → 3.7 Security: keys, data, prompt injection
  • Deeper optimization → 4.7 Performance and latency
  • back → handbook index

Last updated 2026-10-05

← PreviousSafety: limits and human-in-the-loopNext →Security: keys, data, prompt injection

© 2026 Siim Liimand · SeoWeb

GitHub/AI Handbook/Tallinn, Estonia

59.4370° N, 24.7536° E — Tallinn, Estonia

↑ Top