SeoWeb
  • Work
  • Services
  • CV
  • Contact
    • AI
  1. Home/
  2. AI Handbook/
  3. Agents & Evaluation/
  4. Performance and latency

[ Agents & Evaluation ]

Performance and latency

Target audience: technical + deep-diving non-technical | Prerequisites: 4.6 Production monitoring

What you'll learn

After this document you will be able to:

  • explain what latency is (the wait from request to answer) and when speed matters;
  • name the five causes that slow an answer down and five ways to ease them;
  • assess the real price of every speedup — what the speed was bought with;
  • measure latency both as an average and as the 95th percentile (p95 — 95% of cases are faster than this);
  • run every speedup change through evaluation (see 4.5) so quality doesn't drop.

In plain terms

Speed is not a property but a requirement whose limit depends on who is waiting: a customer in a chat window waits seconds; an overnight email can wait minutes. There are five ways to speed things up — but none is free: each is bought at the expense of quality or system simplicity. That's why you measure before and after and decide by the numbers, not impressions.

Latency: when speed matters

Latency is the time from sending the input to receiving the answer — one of the six core metrics 4.5 already named. Here we look at when it matters.

The tolerable limit is not a technical constant — it depends on whether someone stands waiting for the answer or does something else meanwhile:

Use caseTolerable latencyWhy
Chat with a customera few seconds (2–5 s)the customer waits at the screen — waits a little, then closes the window
A draft for an employee (a proposal, a summary)10–30 sthe employee can do other things meanwhile
Overnight email classificationminutesthe work must be done by morning — in the middle of the night speed doesn't matter
A weekly or monthly reportminutes–hoursnobody stands waiting — only the agreed completion time counts

The rule is simple: the more directly a human waits for the answer, the shorter the latency must be — eight seconds in a chat is a catastrophe, invisibly fast in overnight classification.

Latency is tracked with two metrics:

  • The average says how fast a typical request is — but hides the extremes: a single 30-second case disappears, invisible in the average.
  • The p95 says how bad the experience is for the unluckiest user: a p95 of 20 s means every twentieth customer waits 20 seconds or longer — something the average cannot show.

What slows an answer down

Time accumulates from five places. Before speeding anything up, work out which of them matter in your system — otherwise you speed up the right thing from the wrong end.

  1. A big model. A more powerful model answers more thoroughly but more slowly — the 1.1 picture: a specialist thinks thoroughly, an intern answers immediately. For a simple step the big model is simply slow.
  2. Long input. The more tokens (a token being a piece of a word) travel with the request, the longer reading them takes. The whole conversation history filling the context window (the amount of text the model sees at once; see 3.2) grows not only costs but also the wait — the 2.4 principle “give only what's needed” applies to speed too.
  3. Long output. The answer is written token by token — every extra word takes its own time. Three numbers in structured output (2.2) are several times faster than a three-paragraph summary; trim the output nobody reads.
  4. An agent's many turns. As 4.1 showed, an agent decides itself how many steps it takes — every turn is a new request with its own wait, and ten turns means ten waits added together.
  5. Provider load. Now and then the provider's service slows down or some request fails. You can't influence it, but you must account for it: retries and fallback behavior (3.4) lengthen the wait further.

In plain terms: an answer is produced in three activities — the model reads the input, thinks, and writes the output. Long input is a long read-aloud, a big model is a slow thinker, and long output is long writing. An agent does this trio several times in a row — and the time adds up.

Five ways to speed up

ToolWhat it doesAt whose expense
A smaller or faster model for simple stepsclassification and other simple steps run noticeably faster and cheaperquality risk — every model change goes through evaluation (4.5)
A shorter prompt and outputfewer tokens to read and write (see 2.4 and 2.2)terser wording — check that nothing important got trimmed
A cachesame input → the stored answer comes back immediatelythe answer can go stale — the cache must be refreshed
Streamingthe user sees the answer while it is still being writtentotal time doesn't change — the benefit is only the user's perception
Parallelism for independent workseveral independent requests at once, not in sequencea more complex system build — the results must be assembled again

Three tools deserve a couple more words.

A cache (a stored answer for the same input). When a frequently repeated question always gives the same answer, there's no reason to ask the model again: the first answer is stored, the next ones get it immediately — this is how a frequently-asked-questions page is often solved. Bonus: there is no billable request, so the cache saves costs too — the 3.6 principles work twice here.

Streaming (displaying the answer while it is being written). The answer isn't shown all at once but written before the user's eyes: the first words appear within seconds, the rest follow. Total time stays the same, but the experience changes — the system “answers right away” instead of staying silent and then dumping the whole answer at once. Almost indispensable for chat, useless for background work.

Parallelism. When a task has parts that don't depend on each other (for example the summary of the same email and a language check), they are sent out at once: the wait is then the time of the longest part, not the sum of all.

In plain terms: the cache is a notebook of “already answered” — for the same question the answer is read out, not computed again. Streaming doesn't make the system faster, but it feels that way to the user: the first words say “I'm answering” while the rest follow. What doesn't depend on each other is done at once, not one after another.

The triangle: speed, quality, cost

Every speedup is a purchase made at the expense of some other value. Three values form a triangle whose sides cannot all be stretched to maximum at once:

  • you want it faster → either you pay more (a faster service tier, a cache — cost rises) or you take a risk (a smaller model, shorter instructions — quality may drop);
  • you want it better → a bigger model or a more thorough input — time and cost grow;
  • you want it cheaper → a smaller model, shorter input and output, a cache — you save time at the expense of certainty.

All three cannot be maximal at the same time — every use case chooses its own emphasis: in chat speed comes first (the customer waits), in overnight batch work cost (nobody waits, but the bill arrives).

And one firm rule: every speedup attempt goes through evaluation — with metrics before and after (4.5). If time was cut in half and quality stayed the same, the change is good; if quality dropped, it is rolled back — speed bought at the expense of customer dissatisfaction is not a purchase. Cost savings and speed savings go hand in hand: the same means lower both (3.6).

After launch, production monitoring (4.6) watches the numbers. And when the system grows into multiple workflows and agents, keeping the triangle becomes an architecture decision — 5.1 covers that.

A step-by-step example: the CallHelp chatbot

The same CallHelp phone support where 2.4 cleaned transcripts and 4.3 built the customer profile added a chatbot to its site: the customer writes a question, the system puts a big model and the whole conversation history into the request (following the 3.2 principle), and the model answers. The answer took 8–12 seconds — customers closed the window before the answer.

Step 1 — measurement. One week of data: average latency 9.5 s, worst case (p95) ~20 s — every twentieth customer waited at least 20 seconds. Most of the time went to reading the long input (~4,000 tokens of history on every request) and writing the long answer with the big model. Numbers on paper first, changes after.

Step 2 — three changes.

  1. Classification on a small model (0.8 s). Every question first goes through a small, fast model that determines the topic (order status, return, price list) and picks the needed customer data. For this simple step the big model isn't needed (1.1: the intern) — 0.8 s and the step is done.
  2. A customer profile instead of the long history. The customer is identified and the system puts only the structured profile into the request — the “state of affairs” (4.3): name, orders, past resolutions. The input is ~400 tokens instead of ~4,000 tokens of history — the same facts, a tenth of the reading.
  3. Streaming in the chat. The answer reaches the customer token by token: the first words appear at around 1 s, although the full answer completes in a few seconds, as before.

Step 3 — evaluation and result. Before launch the updated system went through the test set (4.5): quality did not drop — the profile contained the same facts as the long history. A month later:

MetricBeforeAfter
Average latency9.5 s2.8 s
p95~20 s~5 s
First words to the customer9.5 s (the whole answer at once)~1 s
Cost per daybaseline−35%
Quality (test set, 4.5)baselinedid not drop

The average time shrank more than threefold, the worst case almost fourfold — and customers now stay to wait for the answer. Cost fell 35% thanks to shorter input and a smaller model: cost savings and speed savings walked hand in hand.

Summary

  • Latency is the wait from input to answer; the tolerable limit depends on whether anyone waits — a few seconds in chat, minutes in background work.
  • Time accumulates from five places: a big model, long input, long output, an agent's many turns, and the provider's momentary load.
  • Five accelerators: a smaller model for simple steps, a shorter prompt and output, a cache, streaming, and parallelism — each bought at the expense of something.
  • The triangle: speed, quality, and cost cannot all be maximal at once; every use case picks its emphasis and every change goes through evaluation (4.5).
  • Measure the average and the worst case (p95): the average tells the typical experience, p95 shows the unluckiest user's — the average hides the extremes.

What's next?

  • previous → 4.6 Production monitoring
  • next level → 5.1 Architecture at scale (level 4 is complete)
  • back → handbook index

Last updated 2026-10-05

← PreviousProduction monitoringNext →Architecture at scale

© 2026 Siim Liimand · SeoWeb

GitHub/AI Handbook/Tallinn, Estonia

59.4370° N, 24.7536° E — Tallinn, Estonia

↑ Top