- Home
- AI Handbook
- Agents & Evaluation
[ Level 04 ]
Agents & Evaluation
Agent systems, RAG, long-term memory, multi-agent architectures, evaluation, monitoring, and latency.
Agent systems: what they are and when you need one
What makes an agent an agent — goal, tools, and decision-making freedom — when to pick a workflow instead, and why a hybrid with one agent step is often the best choice.
Updated 2026-10-05
RAG: using your own data as a source of answers
The three ways to give a model your own knowledge, the four-step RAG flow from chunking to a sourced answer, and why a RAG system needs ongoing maintenance.
Updated 2026-10-05
Long-term memory and state management
How to give an AI system memory across sessions — what to store, how to update and forget it, how to keep state, and the two risks: stale data and context poisoning.
Updated 2026-10-05
Multi-agent architectures
When one agent is no longer enough: the three multi-agent patterns, why a simple workflow usually does the orchestrating, and how costs, errors, and testing multiply.
Updated 2026-10-05
Evaluation: how to know whether a system is good
Why “looks good” is not a metric: build a test set and a golden set, track six core metrics, and catch regressions before a change reaches your customers.
Updated 2026-10-05
Production monitoring
Monitoring versus evaluation: six metrics to watch, specific alerts with thresholds, and a 15-minute weekly routine that feeds your test set and prompt updates.
Updated 2026-10-05
Performance and latency
When speed matters, the five things that slow a response down, five ways to speed it up, the speed–quality–cost triangle, and how to measure p95.
Updated 2026-10-05