SeoWeb
  • Work
  • Services
  • CV
  • Contact
    • AI
  1. Home/
  2. AI Handbook/
  3. Enterprise Scale/
  4. Architecture at scale

[ Enterprise Scale ]

Architecture at scale

Target audience: technical + non-technical leadership | Prerequisites: 4.7 Performance and latency

What you'll learn

After this document, you'll be able to:

  • name what multiplies as you grow: systems, people, keys, and logs;
  • explain the four architecture patterns and say for each one when it pays off;
  • grasp the question specific to AI: many systems use the same model — one model change (5.5) or a price change affects all of them at once;
  • maintain an architecture sheet — a single page with all AI components and their dependencies;
  • ask the scaling question “what happens when users ×10?” and know what to look at when answering it.

In plain terms

At the end of 3.1 API integrations there was one system calling one model — and it worked. At scale, a different question gets asked: not “does the system work?”, but “can anyone keep the systems running when there are a dozen of them and several people touch them?”. Architecture is a set of decisions: how the work is divided into pieces, how the pieces talk, and where everything that costs money flows through.

What changes when the system grows

Growth isn't just “more of the same”: four things multiply and change character.

In one systemAt scale
Systemsone return flowmultiple services (a service is an independently maintainable part of a system): classifier, search, shipping — each lives its own life
Peopleone developer who knows everything by hearta team — nobody knows everything by heart anymore; knowledge has to be written down
Keysone API key (the system's passport and cash register)every service has its own key — and the question of where to review them
Logsone log, read by one personmany logs — you need an understandable overview, not ten files

Each row raises a question — and the four patterns below are the answers to them.

In plain terms: growth doesn't add details — it multiplies them. Two systems mean two keys, two logs, two places where a bug can hide. That's why the recipe at scale isn't “it'll work with more too,” but a routine: split into pieces, let the pieces talk in an agreed way, and put everything that costs money behind one place.

The four architecture patterns

PatternWhat it isWhen it helps
Separation into servicesthe system is split into services: each function is its own part — classifier, search, shipping — maintained separatelywhen parts are updated at different rhythms or by different people; when one part's bug must not stop everything
Event-driven design (event-driven)services talk through events (an event is a message that triggers an action): “a new email arrived” triggers the flow — nobody calls anyone directlywhen there are, or will be, more senders and listeners: a new service can start listening to an event without changing the existing ones
The central access gateway (a gateway — the central point for all model calls)all model calls go through one place: cost visibility, key management, and limit control in one spotwhenever more than one service calls the model — at scale, practically always
Environment separationenvironments (dev/test/production): experiments in the dev and test environments; production serves real userswhenever there is something to protect: you don't experiment with production data and you don't test with the production key

In plain terms: the patterns' ideas are everyday ones. Separation into services is splitting a company into departments — everyone does their own thing, and if one gets sick, the whole house doesn't stand still. Event-driven design is a notice board: the sender pins up a notice, and a new department reads the same notices along the way. The access gateway — the central gateway for short — is a shared cash register where you can see all spending. Environment separation is a fitting room: what's meant for sale is tried on somewhere else first.

You don't adopt all the patterns at once — each one costs something of its own: splitting into services brings network latency, events bring queues and delays, and the gateway is one more part to maintain. A small system kept running by one person needs none of them; the first sign that the time has come is a simple sentence from a developer: “I honestly no longer know what runs where.”

Whether the architecture holds is shown by the scaling test. Scaling (scaling) means accounting for growth before it happens: “what happens when users ×10?” The answer lives in three places:

  • Limits. Can the provider's onboarding limit carry ten times the call volume — and can one careless loop take capacity away from the others (3.1 on limits)?
  • Costs. Has anyone seen and approved the tenfold bill in advance (3.6 Cost management; for the big picture, 5.4)?
  • Latency. Does the wait time stay tolerable at peak load — on average and in the worst case (4.7)?

If any answer is “I don't know,” the test found the weak spot before a user did.

One model, many systems: the central gateway

AI has one particularly sharp question at scale: many systems use the same model. When the provider announces a model change (5.5) or raises a price, it affects not one system but all of them at once. Without organization, this is discovered as the systems stop working one after another.

The answer is a central gateway and written-down dependencies:

Without a gateway:                     With a gateway:

classifier ──────► model               classifier ──┐
chatbot ────────► model                chatbot      ├─► central gateway ─► model
weekly report ───► model               weekly report┘        │
                                                             └─► one log: who, which one, how much

three keys, three places to change     one key, one place to change

Next to the gateway goes a list of dependencies — which system uses which model and why:

ServiceModelWhy this one
The email classifiersmalla simple step — speed and cost matter (4.7)
The chatbotlargethe customer is waiting — quality matters
The weekly reportlargeonce a month — speed doesn't matter

When a model change or price change looms, the list is the manager's answer to the question “who is affected?” and the technician's work plan: with a gateway, the change is a one-place affair; without one, it's a repeated expedition from system to system.

In plain terms: the central gateway is a shared payment system and switchboard in one: everything that costs money passes through one register — and one switch reroutes the whole house. But a switchboard only helps if it's written down which cable goes where; that's what the dependency list is for.

The architecture sheet: one page that has to be right

The architecture sheet (a one-page diagram of systems and dependencies) is one page holding all the AI components and their dependencies: which services exist, what triggers what, what goes through the gateway, which model everyone uses.

HOMECRAFT STORE — AI SYSTEMS (architecture sheet, updated 2026-10)

webshop ──("new email")──► return service ──► central gateway ──► models (see the dependency table)
                              │
                              └──► draft ► Piret approves ► shipping
Keys: at the gateway | Logs: the gateway's shared log | Environments: dev / test / production
System owner: Tanel

Three rules that make the sheet right:

  • one page — if it doesn't fit, the system is already too complicated to change;
  • updated on the day of the change — an old diagram is worse than a missing one, because it eats trust;
  • owner written down — next to every system, one name: responsibility always belongs to a person, not to a system (1.6 Roles and responsibility).

The reason for the page is human: when the developer leaves, the next person can figure it out.

Step-by-step example: the HomeCraft Store becomes a chain

The starting point. Mari's HomeCraft Store — Tanel develops, Piret approves drafts — has grown into a chain of 12 stores. The return flow that was built in 2.5 and made into a system in 3.1 worked perfectly in a single store.

1. How it got there. For every new store, Tanel copied the flow: its own key, its own log, the same prompt. Fast — nothing new was built.

2. What broke.

  • 12 keys. The question “under whose key does this cost run?” went unanswered.
  • 12 logs. A customer email from store 7 went without a draft; the reason was found only after reading through ten files.
  • One model change broke ALL of them. The provider announced the end of support for the old model. Tanel fixed one flow, then another — for three days, some stores had no drafts and Piret did the work by hand.

3. The solution: one central return service.

  • One service, not twelve clones: one classifier, one draft planner, one API key.
  • The stores call on an event: “a new email arrived at store 7” — a store doesn't need to know what happens behind the scenes; adding a new store is a new sender, not a new system.
  • All model calls through the gateway: cost and limits visible from one place.
  • Environments separate: Tanel tries a new prompt in the dev environment without a real customer; production serves all 12 at once.
  • The architecture sheet on one page: services, models, keys always at hand; the gateway's shared log became, as a side effect, the single source for monitoring (4.6).

4. The result.

Before: 12 clonesAfter: one service
API keys121 (at the gateway)
Model change12 places, days1 place, hours
“Who spent?”opinionexactly, from the gateway's log
New storea new copy of the flowone entry — the service already exists

In plain terms: Mari doesn't know what happens inside the gateway, and she doesn't need to — she only knows that the question “what happens when the model changes?” is now one decision, not twelve repairs. If Tanel leaves tomorrow, the next person reads one page and carries on.

Summary

  • Growth multiplies: one system → many services, one developer → a team, one API key → several, one log → the need for an overview.
  • Four patterns answer four questions: separation into services (how to divide into pieces), event-driven design (how things talk), the central access gateway (where everything that costs money flows through), environment separation (where you experiment).
  • One model, many systems: a model change (5.5) or a price change affects all of them at once — a central gateway and written-down dependencies make the change a one-place affair.
  • The architecture sheet is one page, always right, with an owner: when the developer leaves, the next person can figure it out.
  • The scaling test — “users ×10?” — checks the limits (3.1), the costs (3.6, 5.4), and the latency (4.7); “I don't know” is an answer better heard before users do.
  • Agreements about who may make changes and how they get approved are the topic of 5.2 Team workflows and standards.

What's next?

  • previous → 4.7 Performance and latency (level 4 is complete)
  • next → 5.2 Team workflows and standards
  • back → handbook index

Last updated 2026-10-05

Next →Team workflows and standards

© 2026 Siim Liimand · SeoWeb

GitHub/AI Handbook/Tallinn, Estonia

59.4370° N, 24.7536° E — Tallinn, Estonia

↑ Top