SeoWeb
  • Work
  • Services
  • CV
  • Contact
    • AI
  1. Home/
  2. AI Handbook/
  3. Enterprise Scale/
  4. Security and data protection (GDPR, audit)

[ Enterprise Scale ]

Security and data protection (GDPR, audit)

Target audience: non-technical leadership + technical | Prerequisites: 5.2 Team workflows and standards

NOTE TO THE READER: This document is an educational overview, NOT legal advice. It teaches which questions to ask and which processes to put in place — it doesn't replace the law or an expert. Before launching a real system, consult a data protection expert.

What you'll learn

After this document, you'll be able to:

  • say what GDPR (the European Union's General Data Protection Regulation) is and what counts as personal data (data that makes a person identifiable) in an AI system;
  • name the five core requirements and show where each one lives in your system;
  • ask an AI provider five questions about data and know that the responsibility stays with the company;
  • put in place a four-step response plan for a security incident (an incident — data loss or a leak);
  • start preparing for an audit (auditability — being prepared for scrutiny): inventory, trail, records.

In plain terms

Until now, the handbook has looked at security through a technician's eyes — keys, data volume, failover (3.7). This document looks at the same house through the law's window: what is mandatory and who gets asked when something goes wrong. The principle: all data that leaves the system is your responsibility even when it travels to another company for processing. The provider supplies the labor, not the responsibility. Responsibility can be carried only if you know what data you have, where it is, and how long it's kept — that's what an inventory and an audit trail are for, not good intentions.

Personal data and GDPR in an AI system

GDPR is the body of rules about how people's data may be collected, kept, and used — and it applies to your AI system too, if it touches customers. The first check is often the whole answer: does the system contain personal data at all?

In an AI system's context, personal data typically includes: the customer's name, order history, the content of emails and chats. Even a customer number can be tied to a person. If a customer mentions their health or financial situation in a chat, the data is even more sensitive — as we saw in the e-pharmacy example in 3.5.

The five core requirements:

ObligationWhat it meansWhere it lives in your system
Notification (transparency)the customer knows that data is processed, for what purpose, and to whom it's passedprivacy terms + the “an AI provides the answer” label
Data minimization (only what's necessary)only what the task needs goes to the provider and into memoryinput preparation (2.4, 3.7)
Storage limit (how long to keep)every data type has a lifetime and automatic deletion written downthe settings for logs, memory, and backups
Right to be forgotten (the ability to delete)“forget me” must be fulfilled everywhere: in the database, in memory, in logsthe delete action (4.3)
Secure processingaccess only by role; keys out of the draweraccess rights and keys (3.7)

For the non-technical manager: these five rows are the answer to “what is mandatory?”. For each row, ask whether it's in place with us and whose name answers for it — on roles, see 1.6.

Data to an external provider: the contract and responsibility

When the system sends data to an external AI provider, it isn't “just an API call” — legally it's handing data to a data processor (a processor — one who processes data on the company's behalf, e.g., an AI provider). The decision about which data goes out was made by your company — and the responsibility stays there (1.6).

Five questions to ask about every AI provider:

#QuestionWhy it matters
1Is there a data processing agreement?a contract that leaves the provider only the intended use
2Where are the servers located (which country)?additional conditions depend on the country — interpret with an expert
3Are requests used to train the model?with sensitive data, choose the setting where they aren't (3.7)
4How long does the provider keep logs, and who can see them?the provider's log is also a store of your customers' data
5How is data deleted — including at subcontractors?the deletion obligation must hold across the whole chain

Non-compliance is your responsibility: if the provider uses data for training although that wasn't agreed, or loses the data, the customer and the supervisory authority turn to you — the answer “it was the provider's fault” doesn't convince, because you chose the provider.

In plain terms: a data processing agreement is like the handover sheet when your car goes in for repair: it states who receives the car, what may be repaired, and what happens in case of damage. Without the paper, your word stands against theirs — and if the car comes back scratched, your insurance pays.

AI-specific questions: memory, automated decisions, the trail

1. Can other customers' data leak out of the answers? Long-term memory and logs (4.3) are a new risk: if memory doesn't separate customers cleanly, one customer's information can end up in another customer's answer. The protection is in the design: memory and logs are separated by customer, and only someone with role-based rights can see the logs.

2. Does the system make automated decisions about people? The automated decision principle (GDPR Article 22): a person must be able to challenge a purely automated decision — for example, a refusal of a discount — and reach a human. Hence two requirements: in sensitive decisions a human is in the loop (3.5) — the system drafts, the human decides and bears responsibility — and the challenge path is visible to the customer: which human answers, and when.

3. How do you show that the action was correct? The audit trail (a record of every action) is familiar from the protection layers in 3.5: what, when, why. The trail isn't only for investigating incidents — it is evidence: the question “why did the system give exactly this answer?” gets its answer from records, not memories. One limit: the trail is itself a dataset — the log, too, must have a storage limit.

A security incident: when something goes wrong

A security incident (an incident — data loss or a leak) is a question of “when,” not “whether”: a leaking key, a file sent to the wrong address, a provider's notice of a breach. The response has four steps:

StepWhat is doneWho (1.6)
1. What happened?facts written down: which data, whose data, where it went, how much, whenmaintainer
2. Containthe leak stopped: access revoked, keys rotated, if necessary the system taken down (3.5 stop button)maintainer + builder
3. Notifywhere applicable, to the supervisory authority (in Estonia, the Data Protection Inspectorate) within 72 hours; with higher risk, also to customerssponsor
4. Learnwhat happened, why, what changes — written down and verified that the change was madethe whole team

The phrase “where applicable” means that not every incident requires notifications — but the deciding itself takes time; that's exactly why the plan must be ready before the incident: under stress, nobody debates who calls whom.

In plain terms: a response plan is like a fire extinguisher — it's only looked at when something's burning, and that's exactly why it gets hung up before the fire. A four-step plan — explain, contain, notify, learn — must be written so that next to every step stands a name, not “someone.”

Step-by-step example: the Remedy House inventory

3.5 introduced the e-pharmacy chat system “Remedy House”: the system answers questions about stock and prices and routes health questions to a pharmacist. Now the pressure grows: customers' questions increasingly concern medications and health conditions — particularly sensitive data.

1. The data inventory (which data, where, how long). The first step isn't fixing but finding out:

Data typeWhereHow longNeeded by
Customer name and order historycommerce systemas required by law (e.g., invoicing)pharmacist, accounting
Chat text (question + answer)AI provider's logscurrently 2 yearsonly for bug investigation
Health details in answerschat logs—shouldn't reach the log at all
The system's activity trail (what, when, why)our own log system1 yearthe maintainer, for auditing

2. The weak spots the inventory showed — three flaws found, none of them “a hacker's work”:

  • chat logs were kept for 2 years — no storage limit (how long to keep);
  • the pharmacist wrote customers' health details into the answers — data minimization (only what's necessary) wasn't holding;
  • there was no data processing agreement with the AI provider — data outside without a contract.

3. The fixes:

  • chat logs 90 days, then automatic deletion;
  • answer templates without health details: health information goes to the pharmacist in their own system, not into the chat log;
  • the data processing agreement signed with the provider — inside it: no training, a deletion procedure, the server country;
  • text visible to the customer: “This chat is processed by AI; the chat record is kept for 90 days.”

4. An incident drill. The manager asks from above: “If a log file went to the wrong recipient, what now?” The team practices the plan:

StepRemedy House does
What happenedwhich file, whose data, how many records, where it went — facts written down within half a day
Containa written request to the recipient to delete and refrain from further distribution; automatic report sending stopped
Notifya severity assessment: a file with health details — where applicable, to the Data Protection Inspectorate within 72 hours
Learnreports from now on only to the named address, access by roles, the drill written down

In plain terms: the inventory didn't find an attacker — it found three shy flaws that all came from one absence: nobody had written down what happens with the data. And the plan has been practiced: at the first real incident there's no good moment to search for where the fire extinguisher hangs.

Summary

  • Personal data is usually present in an AI system — customer name, orders, the content of emails and chats. The five core requirements (transparency, data minimization, storage limit, right to be forgotten, secure processing) answer the question “what is mandatory.”
  • The provider is a data processor: ask about the agreement, the country, training, logs, deletion — and remember that non-compliance remains the company's responsibility (1.6).
  • AI adds three questions: memory and logs must not mix customers (4.3); in an automated decision about a person there must be a way to challenge it — human-in-the-loop (3.5) and a visible challenge path; the trail makes correctness demonstrable.
  • A security incident (an incident): what happened → contain → notify (where applicable, within 72 hours) → learn. The plan written down and practiced before the incident.
  • An audit (auditability — being prepared for scrutiny) is the sum of three things: inventory, trail, records. When a supervisory authority or a customer asks “how do you process data?”, the answer must be written down — and true.
  • The technical protection layers remain in the neighboring documents: 3.5 and 3.7.

What's next?

  • previous → 5.2 Team workflows and standards
  • next → 5.4 Cost strategy at scale
  • back → handbook index

Last updated 2026-10-05

← PreviousTeam workflows and standardsNext →Cost strategy at scale

© 2026 Siim Liimand · SeoWeb

GitHub/AI Handbook/Tallinn, Estonia

59.4370° N, 24.7536° E — Tallinn, Estonia

↑ Top