SeoWeb
  • Work
  • Services
  • CV
  • Contact
    • AI
  1. Home/
  2. AI Handbook/
  3. Agents & Evaluation/
  4. Multi-agent architectures

[ Agents & Evaluation ]

Multi-agent architectures

Target audience: technical + deep-diving non-technical | Prerequisites: 4.1 Agent systems

What you'll learn

After this document you will be able to:

  • decide whether a task needs a single AI agent or a multi-agent architecture (a system where a task's work is divided among several specialized agents);
  • name the three signs that one agent is no longer enough;
  • distinguish and choose between the three core patterns: the sequential chain, the lead + workers, and mutual checkers;
  • explain why the agents' work is often directed by a simple workflow, not another agent;
  • see how costs, errors, and testing effort multiply with the number of agents, and install limits for each agent separately.

In plain terms

Several agents are not “more power” — they are more complexity. Most jobs are done by one agent with good instructions. Several agents come into play when the work is truly divisible and the result needs independent checking. A multi-agent architecture is the solution to a complex task, not a goal — the simple option always stays the preferred one.

When one agent is not enough

Document 4.1 gave the agent basics: three parts, agent vs workflow, the hybrid. But what if even one agent with good instructions can't cope alone? Three signs:

  1. The task's parts need different specializations. If one part of the work requires different instructions, different tools, or even a different model, one agent's instructions cannot say both well. The analyst's instruction says “be brief and precise”, the writer's “keep the story flowing” — in one instruction they get in each other's way.
  2. The result needs independent checking — the writer is not the checker. Whoever wrote a text is a poor critic of their own text — they read what they meant to say. The same applies to an agent. A separate checking agent sees the result with fresh eyes and goes down its own checklist — not the writer's self-satisfaction.
  3. The work is naturally divisible. Five posts are five independent pieces: they can be done separately, in parallel if needed, and each checked separately. Here there is no single giant task, but a pile of small ones.

One checking question before deciding: do the parts need a different instruction or just a different input? If the same agent can, with the same instruction, do five different inputs five times, five agents aren't needed — one agent with five calls is enough.

The three core patterns

PatternHow it worksWhen it fits
Sequential chain (pipeline: each agent does its part and passes the result on, like an assembly line)agent 1 → agent 2 → agent 3; each one's output is the next agent's inputwhen the work is naturally sequential and the stages are visible already on paper: topics → text → check
Lead + workers (orchestrator + workers: the lead agent divides the work, calls the workers, and collects the results)the lead decides who does what; the workers do their part and hand the results back to the leadwhen the number and shape of subtasks isn't known in advance and “who does what” is itself a decision that requires thinking
Mutual checkers (generator-critic: one does, the other critiques, the doer fixes)work → critique → fix → new critique, until the checker is satisfiedwhen quality is critical and an error is expensive: facts, reports, text going to a customer
(a) SEQUENTIAL CHAIN (pipeline)

[agent 1] ──▶ [agent 2] ──▶ [agent 3] ──▶ result
  topics         texts         check
each agent does its part and passes the result on

(b) LEAD + WORKERS (orchestrator + workers)

              [lead agent]
                    │  divides the work, collects the results
      ┌─────────────┼─────────────┐
      ▼             ▼             ▼
  [worker]      [worker]      [worker]

(c) MUTUAL CHECKERS (generator-critic)

[does] ────────── the work ──────────▶ [critiques]
  ▲                                   │
  └──────── fix instruction ◀─────────┘
      repeats until the checker is satisfied

In plain terms: the three patterns are like a kitchen. The sequential chain is the assembly line: soup → main course → dessert. Lead + workers is the head chef who divides the orders among the cooks and assembles the plate. Mutual checkers are the cook and the taster — the plate doesn't go to the customer until the taster approves it.

Orchestration: who leads

Orchestration (dividing the agents' work, directing it, and summarizing the results) is the leader's job. The first question is not “how to connect the agents” but “who leads”.

The temptation is to build a lead agent: let one agent decide who does what. But then there is again a model in the leader's position — its decisions are unpredictable and its errors propagate to everything below. Here the 2.3 principle works: often a simple workflow directs the agents. A workflow as orchestrator is:

  • predictable — the same steps every time; if an error comes, you know where to look;
  • cheap — the directing itself costs zero tokens;
  • controllable — you decide which agents launch when, not the model.

A lead agent is justified only when the division of the work itself requires thinking — “who does what” is itself a decision. Even so, put limits on the lead agent too: a maximum number of steps, a fixed result form, and a list of tools (the function calling basics, 3.3).

Costs, risks, and limits multiply

In plain terms: one agent is one paid assistant; two agents are two assistants who must agree with each other — costs and opportunities for error grow faster than the benefit.

In a multi-agent system everything multiplies — the good, but above all the bad:

  • Costs multiply. Every agent has its own requests: its own instructions, its own tool calls, its own thinking tokens. Three agents are not three times the price of one agent — they are three times everything, plus the information moving between them (3.6 Cost management). The checker's “fix and try again” cycle is the most expensive part if you don't cap the number of rounds.
  • Errors propagate along the chain. The next agent processes the previous one's result trustingly. If the first agent gives a wrong topic, the second writes a convincing text about it and the third reviews a wrong input (3.4 Errors and error handling).
  • Clarity is lost. In a one-agent system the error is in the instructions, the input, or the model's answer. In a five-agent system there are five instructions, five inputs, and the transfers between them — if the end result is bad, it is considerably harder to say where the error happened.
  • Testing gets complicated. A one-agent test is “input in, output out”. A multi-agent system's test needs combinations: what happens if one agent gives an unexpected output, and how does the next one behave? How to measure all this is covered by 4.5 Evaluation.
  • Limits for each agent separately. A safeguard (a restriction built into the system's design) applies to each agent separately, not to the system at once: the topic finder gets read-only access, the composer has no publishing capability, the check agent has no sending rights. If one agent has all the rights, all the other limits are just words (3.5 Safety).
  • The human stays in the loop for final decisions. Five agents that agree are not five signatures: the agents' consensus is not human approval. Anything going to the customer, involving money, or touching the public always passes through a human.

A step-by-step example: WordMeadow's weekly content

The fictional “WordMeadow” — a content marketing agency. The task: for a client, every week 5 social media posts and 1 newsletter. The work is divisible (topics → text → check), the result needs independent checking (the texts go to the client's public channel), and the rhythm is fixed. Conclusion: three agents, with a simple workflow as the lead.

WORDMEADOW'S WEEKLY CONTENT — the lead is a SIMPLE WORKFLOW (fixed steps, not a lead agent)

 MONDAY                    WEDNESDAY                 THURSDAY                 FRIDAY
─────────────────────────────────────────────────────────────────────────────────────
 [1] topic finder    ──▶   [2] text composer   ──▶   [3] check agent    ──▶   HUMAN
     tool:                 tool:                    checklist:             approves
     the client's past     the brand                facts | style |        and publishes
     posts + analytics     style guide              forbidden claims
                                                        │
                                                        ▼ if a point fails
                                               fix instruction back to [2]
AgentGoalToolsLimits
Topic finder5 post topics + 1 newsletter topic, each with a short justification (why this one)the client's past posts and view analytics — read-onlywrites no text; doesn't justify anything it can't find in the data; output in a fixed form
Text composertexts according to the style guidethe brand's style guideinvents no facts or prices — uses only the given topic and the list of facts; cannot publish anything
Check agentreview of every text against the checklistthe checklist + the client's fact list (prices, dates)doesn't fix anything itself — writes a fix instruction to the composer; has no publishing rights

The check agent's checklist:

CheckpointWhat is checkedExample error
Factsevery number, price, date, and name matches the client's fact lista post names a price that isn't in the price list
Stylethe style guide's voice, length, and vocabularya salesy tone where the guide calls for advisory
Forbidden claimspromises the client must not make“we guarantee results in 30 days”

What happened without a checker. Before the check agent was added, the chain ran straight: topics → texts → the human approves. On Friday the human approved five texts at a glance — and one post went out with the price “29 €” that isn't in the client's price list: the composer had filled the knowledge gap. The client's followers started pointing at the “discount price” and WordMeadow's trust wobbled. A propagated error is born in one agent, but is paid for on all the following steps — the check agent would have compared the price with the price list already on Thursday, not the client on Friday.

Today the workflow doesn't move forward on Thursday until the checklist is fully passed — and on Friday the human has the last word: the agent never publishes anything by itself.

Summary

  • Several agents come into play for three reasons: the parts need different specializations, the result needs independent checking (writer ≠ checker), the work is naturally divisible.
  • Three patterns: the sequential chain (assembly line), lead + workers (the lead divides and collects), mutual checkers (does and critiques, until satisfied).
  • Often a simple workflow directs the agents, not a lead agent — predictable, cheap, and clear when errors need finding.
  • Costs, errors, lost clarity, and testing effort multiply with the number of agents; safeguards apply to each agent separately and the human stays in the loop for final decisions.
  • The principle: a multi-agent architecture is the solution to a complex task, not a goal. Start with one agent or a workflow and add an agent only when you have a concrete reason.

What's next?

  • previous → 4.3 Long-term memory and state management
  • next → 4.5 Evaluation: how to know whether a system is good
  • Agent basics and agent vs workflow → 4.1 Agent systems
  • Safeguards and human-in-the-loop → 3.5 Safety
  • back → handbook index

Last updated 2026-10-05

← PreviousLong-term memory and state managementNext →Evaluation: how to know whether a system is good

© 2026 Siim Liimand · SeoWeb

GitHub/AI Handbook/Tallinn, Estonia

59.4370° N, 24.7536° E — Tallinn, Estonia

↑ Top