SeoWeb
  • Work
  • Services
  • CV
  • Contact
    • AI
  1. Home/
  2. AI Handbook/
  3. System Architecture/
  4. API integrations: calling the model from your program

[ System Architecture ]

API integrations: calling the model from your program

Target audience: technical + non-technical readers who want to go deeper | Prerequisites: 2.6 Prompt management as an asset

What you'll learn

After this document you will be able to:

  • explain what an API (a programming interface) is and when a workflow is no longer enough;
  • handle an API key — your system's passport and cash register — responsibly;
  • assemble a request: the API endpoint, the model name, the prompt (the instruction given to the model), and the parameters, including temperature;
  • read the structure of a response: where the model's text is and how many tokens were spent;
  • move the 2.5 return flow onto the API and test the first call before production.

In plain terms

At level 1 you talked to the AI model yourself: you copied the message in and read the answer. The API means your system does that same work: a program sends the model a message in a fixed form — instructions and data — and receives an answer in a fixed form. The 2.5 workflow already did this, but hidden behind clicks: a tool sent the requests on your behalf. Once you see that five things go out — the endpoint, the model name, the instructions, the parameters, and the key (the last one in a header) — and that the text and the cost are read back from the response, the API logic is clear.

What an API is and why call the model through it

An API (a programming interface) is an agreement on how programs talk to each other: to which address, in what form to send, and in what form the answer comes back. A fitting analogy: a service window. You don't go into the kitchen: you leave your order at the window and get your food back in a set form. The kitchen may be in another city — the window protocol stays the same. The model's "kitchen" sits on the provider's servers; your system only needs the window.

Why not stay with the visual tool? Because 2.5 promised: “when the flow later has to be triggered from inside your own system, the next level is the API”. A tool (n8n, Make, Zapier) is a good place to build, but when the flow must be started by your own program — a new order, a button press in an app — the tool is a middleman to skip: the program calls the model directly. The rule: a clickable process stays at the workflow level; the part that belongs to your program moves to the API.

Here the program calls the model; the reverse direction — the model itself calling tools — is the topic of 3.3 Tools and actions.

The API key: your system's passport and cash register

In plain terms: an API key is a work permit and a cash register in one: it tells the provider who is coming and whose account the costs land on. The key is secret like a bank card PIN — why and how to store it, 3.7 Security teaches.

When your program stands at the window, it has to identify itself. That's what the API key is — a secret code that identifies your system and books the cost to your account. Two consequences:

  • Passport. Everything that calls with your key is counted as your system's doing. That's why every system gets its own key: the store system has one, the test another — if something happens, you see which system caused the error and shut down only that one.
  • Cash register. Every call costs (tokens, see below) and the bill goes to the key's owner.

Three rules: never put the key where others can reach it (open code, an email); replace it immediately if the key has ended up where it must not; delete a key that isn't used. Where to keep a key and how to fix a leak, 3.7 Security teaches.

Request and response: what a call looks like

A request is your system's letter to the model. Four parts:

  1. The API endpoint — where the letter goes; the provider gives it to you in writing.
  2. The model name — which model: providers have many, with different skill and price.
  3. The input instructions — the prompt (the instruction given to the model): the template that has passed 2.1's testing cycle, plus this letter's data.
  4. The parameters — the settings you tune the call with. The best known is temperature: how much the model may vary in its answer. At zero the model almost always picks the same, most probable continuation; higher up it allows itself more variation. Low for a classifier, higher for a writer. You can also set a timeout on the request (if the response doesn't arrive within that time, the system stops waiting); what to do then, 3.4 covers.

The common pseudo-shape of these parts (in reality a couple of technical lines are added — the principle stays the same):

{
  "endpoint": "https://api.provider.ee/v1/calls",
  "model": "language-model-v3",
  "temperature": 0.2,
  "instructions": "You are the classifier of e-store customer messages … CUSTOMER_MESSAGE: [message text]",
  "key": "sk-…"
}

(In reality the key usually travels in a separate header (an HTTP header) — same principle: every call carries it along.)

A response comes back in the same spirit:

{
  "call_id": "req_7f31c",
  "text": "{\"type\": \"return\", \"product\": \"wool scarf\", \"order_number\": \"1187\", \"reason\": \"the scarf was too tight\"}",
  "tokens": { "input": 486, "output": 52 }
}

From the response the program reads three things:

  1. The model's text inside the "text" field. Note: the content is JSON again — 2.2 structured output arrives as text that the program parses.
  2. Token usage. A token is a piece of text on the basis of which the model reads and writes; input + output is the basis of the cost — how the token count becomes euros, 3.6 Cost management teaches.
  3. The call ID — every call is an addressable record: later you can see what happened with it (2.5's three-number tracking).

In plain terms: request and response are an exchange of letters in an agreed form: your system writes on its side, the provider on theirs — neither has to know the other's inner life, only the form.

The rate limit: how many times per minute

In plain terms: the rate limit is a bus door: a fixed number of passengers get through at a time. It isn't there to hold you back, but so that everyone with a ticket gets in.

The rate limit is the provider's agreed cap: how many requests per minute (often also per day) one key may send. Why does it exist? The model is served from shared capacity — one loop running in the wrong place would take the capacity away from the others. The limit protects everyone, you included: without it, your system would be left waiting because of someone else's error.

When the limit is exceeded, the request is left unfulfilled and an error message comes back — the system's job is to wait and try again; how to do that properly (backoff, retries, fallback paths), 3.4 Errors and error handling teaches.

The HomeCraft Store's yardstick: 40 messages a day over an 8-hour working day averages about 5 messages an hour — many times below the limit. The limit becomes relevant when the system processes hundreds of messages at once or some loop repeats calls: before scaling up the volume, check the limit.

The first call: testing before production

The first request isn't 40 real customer messages but one test message — the same discipline as 2.5's testing table:

  1. Take the simplest of the test cases: a typical return with a product and a number.
  2. Send one request with the test key and read the response yourself: is "type" "return", are the fields filled in, is the token count reasonable?
  3. Send the same request a second time. Did the response stay the same? (Why this matters, the example below shows.)
  4. Go through another three or four cases — including the "none" and "other" variants.
  5. Only then switch to production — and even there the human-in-the-loop remains: Piret keeps approving the drafts. The API changes who calls the model, not who decides.

For large volumes one rule: a lot at once isn't a test, it's production without a safety net.

Step-by-step example: moving the return flow to the API

Tanel, the HomeCraft Store's developer, moves 2.5's return flow from the visual tool into the store's own system: for every new message, the program assembles the request itself.

1. What stays untouched. The classifier prompt isn't written into the code — the program loads it from the prompt bank as the approved version (v2). When the prompt is improved and approved, the system updates without touching the code — that's exactly why prompts are an asset.

2. The trigger. A new message to the store's address starts the flow as before; only the trigger is now an event in the store's system, not the tool's blocks.

3. The request. The program assembles the five parts — the API endpoint, the model name, 2.5's classifier prompt (the message's text under [CUSTOMER_MESSAGE]), the parameters (temperature 0.2), and the store system's key (the last one in a header):

{
  "model": "language-model-v3",
  "temperature": 0.2,
  "instructions": "You are the classifier of e-store customer messages. Determine type, product, order_number and reason … CUSTOMER_MESSAGE: “Hello! I would like to return the wool scarf (order 1187) — it turned out too tight. Pille K.”",
  "key": "sk-…"
}

4. The response.

{
  "call_id": "req_7f31c",
  "text": "{\"type\": \"return\", \"product\": \"wool scarf\", \"order_number\": \"1187\", \"reason\": \"the scarf was too tight\"}",
  "tokens": { "input": 486, "output": 52 }
}

The program reads the fields: "type" goes into the condition ("other" → Piret's list), "order_number" into the lookup. From there everything is as in 2.5: the draft with the real order's data, Piret confirms before sending.

5. And if the request repeats? Tanel sends the same message one more time — that's testing step 3 from above:

{
  "call_id": "req_7f31d",
  "text": "{\"type\": \"return\", \"product\": \"wool scarf\", \"order_number\": \"1187\", \"reason\": \"the scarf was too tight\"}",
  "tokens": { "input": 486, "output": 52 }
}

The same result. That's exactly what the low temperature guarantees: the same input gives practically the same classification — the classification is repeatable and the errors investigable. The model still gives no hundred-percent guarantee; that's why Piret stays in the chain, and both call IDs stay in the log.

In plain terms: from the outside nothing has changed — the message arrives, the draft gets made, Piret approves. The change is on the inside: instead of clicking blocks, the system sends requests.

Summary

  • An API (a programming interface) is the way a program calls the model: the workflow tool did it hidden behind clicks, with the API your system does it itself — the trigger is your program.
  • The API key is passport and cash register: every system its own key, kept like a bank card, replaced immediately when in doubt.
  • A request = API endpoint + model name + prompt (the instruction given to the model) + parameters; temperature shows how much the model may vary in its answer.
  • From the response structure the program reads three things: the model's text (structured output as text), the token usage (the basis of cost — details in 3.6), and the call ID.
  • Limit, repetition, and testing: the rate limit protects everyone; a low temperature makes the classification repeatable; the first call is made with one test message, not in production — and the human's approval stays.

What's next?

  • previous → 2.6 Prompt management as an asset
  • next → 3.2 Context management: how the model “remembers”
  • Cost details → 3.6 Cost management
  • back → handbook index

Last updated 2026-10-05

Next →Context management: how the model “remembers”

© 2026 Siim Liimand · SeoWeb

GitHub/AI Handbook/Tallinn, Estonia

59.4370° N, 24.7536° E — Tallinn, Estonia

↑ Top