Reference · updated daily

AI Architecture Core

The layers every production AI system is built from, how a request moves through them, and the trade-offs that decide cost, latency and quality. Read top-down for the system view; jump to a layer from the stack.

Daily note

Each day adds one short note on a core architecture idea. Notes appear here once the first one is posted.

Layer 1

Application & experience

Where people meet the system: chat surfaces, copilots inside existing tools, batch jobs, and APIs other services call. This layer owns the contract with the user: what goes in, what comes back, how long they wait, and what happens when the model is wrong.

Components

  • Chat UI, embedded copilots, batch pipelines
  • API gateway, auth, rate limiting, quotas
  • Streaming transport (SSE / WebSocket)
  • Feedback capture (thumbs, edits, escalations)

Design decisions

  • Stream tokens or return a complete answer
  • Synchronous chat vs async job with a callback
  • Human-in-the-loop checkpoints for high-impact actions
  • How citations and confidence are shown
Layer 2

Orchestration & agents

The control logic between the user and the model. It assembles the prompt, decides whether to retrieve, calls tools, loops until a goal is met, and manages what the model remembers. Most product behaviour lives here rather than in the weights.

Components

  • Prompt templates and system instructions
  • Tool / function calling with JSON schemas
  • Planner–executor loops, routers, sub-agents
  • Short-term (context) and long-term (store) memory
  • Tool protocols such as MCP for connecting systems

Patterns

  • Single call: prompt in, answer out
  • Chain: fixed sequence of calls
  • Router: classify, then send to a specialist model or prompt
  • Agent loop: reason → act (tool) → observe → repeat, with a step budget
Agent cost ≈ steps × (prompt tokens + output tokens) — context grows each step, so cost rises faster than linearly without summarisation or caching.
Layer 3

Knowledge & retrieval

Retrieval-augmented generation (RAG) grounds answers in your own documents without retraining. Quality is decided mostly by chunking and ranking, not by the model.

Ingest path

  • Parse (PDF, HTML, tables) and clean
  • Chunk: typically 200–800 tokens with overlap, or by document structure
  • Embed each chunk into a dense vector
  • Index in a vector store (HNSW, IVF-PQ) with metadata filters

Query path

  • Rewrite or expand the query
  • Hybrid search: dense vectors + BM25 keywords
  • Rerank top-k with a cross-encoder
  • Pack the best passages into the prompt with source IDs
similarity(q, d) = cos θ = (q · d) / (‖q‖ ‖d‖)  ·  measure with recall@k, MRR, and answer faithfulness
Layer 4

Foundation model

Almost every current large language model is a decoder-only transformer: a stack of identical blocks that turns a sequence of tokens into a probability distribution over the next token. Variants change the attention pattern, the feed-forward layer, or add encoders for images and audio.

Building blocks

  • Tokenizer (BPE / SentencePiece), vocab ~32k–256k
  • Token embeddings + rotary positions (RoPE)
  • Self-attention: multi-head, grouped-query (GQA)
  • Feed-forward (SwiGLU) or Mixture-of-Experts
  • RMSNorm and residual connections

Variants to know

  • Dense: every parameter used per token
  • MoE: router picks top-k experts; large total, small active params
  • Multimodal: vision/audio encoder projected into token space
  • Long context: RoPE scaling, sliding-window or sparse attention
Inside layer 4

One transformer block

Each block reads the running representation (the residual stream), lets every token look back at earlier tokens through attention, then transforms each position independently with the feed-forward layer. Residual connections (dashed) add each sub-layer's output back to its input, which is what lets very deep stacks train.

Attention(Q, K, V) = softmax( Q·Kᵀ / √dk + mask ) · V
  • Attention mixes information across positions; cost grows with sequence length squared in prefill.
  • Feed-forward holds most parameters (about two-thirds in a dense model) and much of the stored knowledge.
  • GQA shares key/value heads across query heads to shrink the KV cache.
  • MoE swaps the single feed-forward for many experts; only top-k run per token.
Layer 5

Training & adaptation

A model passes through stages, each with different data and objectives. For most enterprise work the choice is not whether to pretrain but which adaptation to apply on top of an existing model.

Stages

  • Pretraining: next-token prediction on trillions of tokens
  • Supervised fine-tuning: instruction–response pairs
  • Preference tuning: RLHF (reward model + PPO) or DPO
  • Distillation: a small model learns from a larger one

Adaptation ladder (cheapest first)

  • Prompting and few-shot examples
  • RAG for knowledge that changes
  • Parameter-efficient tuning: LoRA / QLoRA adapters
  • Full fine-tune for behaviour and format at scale
Training compute ≈ 6 · N · D FLOPs (N params, D tokens)  ·  compute-optimal ≈ 20 tokens per parameter (Chinchilla)
Layer 6

Inference & serving

Serving splits into two phases with different bottlenecks. Prefill processes the whole prompt in parallel and is compute-bound; it sets time-to-first-token. Decode produces one token at a time, rereads all weights and the KV cache each step, and is memory-bandwidth-bound; it sets time-per-output-token.

Techniques

  • KV cache, with paged allocation (PagedAttention)
  • Continuous (in-flight) batching
  • Quantization: FP8, INT8, INT4 weights
  • Speculative decoding with a draft model
  • Prefix / prompt caching for shared system prompts

Metrics

  • TTFT: time to first token (ms)
  • TPOT: time per output token (ms)
  • Throughput: tokens/s per GPU
  • Goodput: requests/s that meet the latency SLO
  • Cost per 1M input / output tokens
Weights memory = params × bytes  →  70B × 2 B (FP16) = 140 GB; at INT4 ≈ 35 GB
KV cache per token = 2 × layers × kv_heads × head_dim × bytes
  e.g. 2 × 32 × 8 × 128 × 2 B = 128 KiB/token → 8,192-token context ≈ 1 GiB per sequence
Inference compute ≈ 2 · N FLOPs per generated token
Inside layer 6

The KV cache, explained

A language model writes one token at a time. To pick each new token, attention compares it with every earlier token using their keys (what each token offers to be matched on) and values (what it contributes once matched). Those keys and values never change once computed, so serving engines keep them in GPU memory instead of recomputing them. That store is the KV cache. It makes generation fast, and it is usually what limits how many people one GPU can serve.

1 · What the cache saves

Without a cache t1 t2 t3 t4 t5 step 1 Step 1, token 1: K and V computed this step Step 1, token 2: not generated yet Step 1, token 3: not generated yet Step 1, token 4: not generated yet Step 1, token 5: not generated yet step 2 Step 2, token 1: K and V computed this step Step 2, token 2: K and V computed this step Step 2, token 3: not generated yet Step 2, token 4: not generated yet Step 2, token 5: not generated yet step 3 Step 3, token 1: K and V computed this step Step 3, token 2: K and V computed this step Step 3, token 3: K and V computed this step Step 3, token 4: not generated yet Step 3, token 5: not generated yet step 4 Step 4, token 1: K and V computed this step Step 4, token 2: K and V computed this step Step 4, token 3: K and V computed this step Step 4, token 4: K and V computed this step Step 4, token 5: not generated yet step 5 Step 5, token 1: K and V computed this step Step 5, token 2: K and V computed this step Step 5, token 3: K and V computed this step Step 5, token 4: K and V computed this step Step 5, token 5: K and V computed this step 15 K/V computations With a KV cache t1 t2 t3 t4 t5 step 1 Step 1, token 1: K and V computed this step Step 1, token 2: not generated yet Step 1, token 3: not generated yet Step 1, token 4: not generated yet Step 1, token 5: not generated yet step 2 Step 2, token 1: K and V read from cache Step 2, token 2: K and V computed this step Step 2, token 3: not generated yet Step 2, token 4: not generated yet Step 2, token 5: not generated yet step 3 Step 3, token 1: K and V read from cache Step 3, token 2: K and V read from cache Step 3, token 3: K and V computed this step Step 3, token 4: not generated yet Step 3, token 5: not generated yet step 4 Step 4, token 1: K and V read from cache Step 4, token 2: K and V read from cache Step 4, token 3: K and V read from cache Step 4, token 4: K and V computed this step Step 4, token 5: not generated yet step 5 Step 5, token 1: K and V read from cache Step 5, token 2: K and V read from cache Step 5, token 3: K and V read from cache Step 5, token 4: K and V read from cache Step 5, token 5: K and V computed this step 5 K/V computations + 10 reads computed this step read from cache not generated yet

To produce token 5, the model needs keys and values for tokens 1 to 5. Without a cache it recomputes all of them at every step. With a cache it computes only the newest one and reads the rest. Over a 1,000-token answer that is 500,500 key/value computations per layer without a cache, against 1,000 with one.

2 · What one token costs in memory

2K and V
×
32layers
×
8KV heads
×
128head dim
×
2 BFP16
=
128 KiBper token

Every token in every active conversation holds this much, in every layer, for as long as the conversation is live. The shape shown is typical of an 8B-parameter model.

3 · How many conversations fit on one 80 GB GPU

weights 16 GBKV cache budget 64 GB (≈ 59.6 GiB)
2K context0.25 GiB per conversation
238
8K context1 GiB per conversation
59
32K context4 GiB per conversation
14
128K context16 GiB per conversation
3

Same GPU, same model. Only the context length changes, and capacity falls from 238 conversations to 3. Long context is a memory decision before it is a model decision.

Assumes an 8B model in FP16 (32 layers, 8 KV heads, head dim 128) and every conversation at full length. Ignores activations and runtime overhead, so real capacity is somewhat lower.

4 · How architects shrink it

Multi-head attention32 KV heads · FP16
512 KiB
Grouped-query attention8 KV heads · FP16
128 KiB
GQA + FP8 cache8 KV heads · 1 byte
64 KiB

Make each token smaller

  • Grouped-query attention: several query heads share one KV head
  • FP8 / INT8 KV cache: halve the bytes per value
  • Latent attention (MLA): store a compressed latent and expand on use

Store fewer tokens, waste less

  • PagedAttention: allocate in small blocks, so little memory sits reserved but unused
  • Prefix caching: a shared system prompt is stored once for all users
  • Sliding-window attention: keep only the most recent N tokens in some layers
  • Offload: move idle conversations to CPU memory or SSD
GPU memory needed ≈ weights + (concurrent users × context tokens × KV bytes per token) + headroom for activations and runtime
Layer 7

Data foundation

Every layer above depends on data that is clean, permissioned and traceable. The same pipelines feed pretraining corpora, fine-tuning sets, retrieval indexes and evaluation sets.

Components

  • Ingestion from lakes, warehouses, SaaS sources
  • Deduplication (exact + MinHash near-dup)
  • Quality and toxicity filtering, PII redaction
  • Labelling and synthetic data generation

Controls

  • Lineage: which data trained or grounded which output
  • Access control carried into the retrieval index
  • Versioned datasets and eval sets
  • Retention and consent rules
Layer 8

Compute & infrastructure

Accelerators, memory and the network between them. For large models, memory capacity and bandwidth matter as much as raw FLOPs, and the interconnect decides how well work splits across devices.

Hardware

  • GPUs / TPUs / custom accelerators
  • HBM capacity and bandwidth per device
  • Intra-node links (NVLink) and inter-node fabric (InfiniBand, RoCE)
  • Managed platforms: Azure AI Foundry, Amazon Bedrock / SageMaker, Vertex AI

Parallelism

  • Data: copies of the model, split the batch
  • Tensor: split each matrix across devices
  • Pipeline: split layers into stages
  • Expert: spread MoE experts across devices
  • Sharded optimizers (ZeRO / FSDP)
Across all layers

Cross-cutting concerns

Evaluation

  • Offline golden sets per use case
  • LLM-as-judge with human calibration
  • RAG: retrieval recall, faithfulness, answer relevance
  • Online A/B tests and regression gates in CI

Safety & guardrails

  • Input checks: prompt injection, PII, policy
  • Output checks: grounding, toxicity, data leakage
  • Tool permissions and least privilege for agents
  • Red-teaming before each release

Observability

  • Traces per request across retrieval, tools, model
  • Token usage and cost attribution
  • Latency (TTFT, TPOT) and error rates
  • Drift in inputs and quality scores

Governance

  • Model and prompt registry with versions
  • Risk tiering and approvals per use case
  • Audit logs and explainability records
  • Regulatory mapping (e.g. EU AI Act, NIST AI RMF)
Across all layers · added from reader feedback

Identity & policy: on whose authority?

Compute and data platforms settled access control years ago with IAM: every call carries a principal, and policy decides what it may touch. AI systems reopen the question, because the caller is increasingly an agent acting for a person. Every layer, from retrieval to tool execution, now has to answer one question: on whose authority is this running, and with what limits?

1 · Authority should narrow at every hop

Usersigned in via SSO
Everything this person is allowed to see and do
↓ delegates a scoped token
Applicationregistered client
Read documents and calendar, on this user's behalf
↓ token exchange for one task
Agentone task, short-lived
Read documents in Project X, for 30 minutes
↓ per-tool scope
Tool callsingle action
Search Project X, read only

The agent never holds more access than the person it acts for, and each step down gets only what that step needs. If a token leaks or a prompt injection hijacks the agent, the damage stays inside the narrowest scope.

2 · The question each layer must answer

LayerQuestionControl
L1 ApplicationWho is the user?SSO with OIDC, MFA, session lifetime
L2 OrchestrationWhat may this agent do for them?Delegated, short-lived tokens (OAuth on-behalf-of, token exchange), step budgets, human approval for high-impact actions
L2 ToolsWhat can this call change?Per-tool scopes and allow-lists, OAuth on MCP servers, read-only by default, dry runs
L3 RetrievalWhich documents may this user see?Source ACLs carried into the index and filtered at query time, before anything reaches the prompt
L6 GatewayWhich tenant, quota and model?API gateway, tenant isolation, rate limits, data-residency routing
L7 DataMay this data be used for this purpose?Purpose-based access, consent flags, lineage
AuditWho did what, for whom?Log the full chain per action: user → agent → tool → resource

3 · Design rules

Enforce outside the modelPrompts can be injected; policy checks in the gateway and tools cannot be talked out of a decision.
Deny by defaultA tool or document is reachable only when a policy explicitly allows it.
Short-lived, scoped credentialsMinutes, not months. No shared service accounts acting for everyone.
Filter before generationTrim retrieval results by permission before the model sees them, never after it answers.
Policy as codeExpress rules in an engine such as OPA or Cedar so they are versioned, tested and reviewed.
Watch for confused deputiesAn agent with broad access can be tricked into using it for a user who lacks it. Carry the user's identity on every call.
Across all layers · added from reader feedback

Governance by design: proving the system does what it should

Governance added at review time produces documents about a system. Governance built in at design time produces evidence of how it behaves. The rule is simple: every intended behaviour becomes a testable requirement, and every requirement leaves a trail that a reviewer or auditor can follow from intent to production.

1 · The traceability chain

  1. Intended useWhat the system is for, who uses it, what it must never do, and its risk tier.
  2. RequirementEach "must" and "must not" written so it can be measured.
  3. Design controlThe mechanism that enforces it: retrieval scope, guardrail, human approval, tool permission.
  4. Eval with a thresholdA test set and pass mark, run in CI. A failed gate blocks the release.
  5. Production monitorThe same check sampled on live traffic, with an alert and an owner.
  6. EvidenceVersioned eval reports, approvals, logs and incidents, produced by the pipeline itself.

Worked example, an HR policy assistant. Intent: answer employee questions from the approved handbook only. Requirement: no answer without a citation to the handbook. Design control: retrieval limited to the handbook index, and the response is rejected if it has no citation. Eval: 300 test questions, at least 95% grounded and zero uncited answers. Monitor: 2% of live answers scored for groundedness each week. Evidence: eval report per release, weekly monitor results, sign-off record.

2 · The evidence pack

EvidenceProduced byWhenWhat frameworks ask for
Intended use & risk tierProduct owner with riskDesignNIST AI RMF Map; EU AI Act risk classification
Requirements as eval setsEngineeringDesign and buildNIST AI RMF Measure; accuracy and robustness obligations
Model, data & prompt cardsEngineeringEvery releaseTechnical documentation; data governance
Eval reports & gate decisionsCI pipelineEvery releaseISO/IEC 42001 performance evaluation
Human-oversight design & approvalsAccountable ownerBefore deployHuman oversight obligations
Logs, monitors & incidentsPlatform teamContinuouslyRecord-keeping; NIST AI RMF Manage
ISO/IEC 42001, the NIST AI RMF and the EU AI Act overlap heavily. Design the evidence once and map it to each, rather than running three separate compliance efforts.

3 · The AI engineering governance skill set

Technical

  • Eval design and metrics
  • RAG, agents and guardrails
  • CI/CD and MLOps pipelines
  • Observability and logging

AI engineering governance

  • Turn policy into testable requirements
  • Set release gates and thresholds
  • Design evidence that is generated, not assembled
  • Explain model behaviour to auditors and the board

Governance

  • Risk assessment and tiering
  • Regulatory mapping
  • Accountability and approvals
  • Audit and assurance

The middle column is the scarce role: someone who can read a model card and a control framework with equal confidence.

How the layers connect

Lifecycle of one request

  1. Receive and authenticateGateway checks identity, quota and tenant (L1).
    I/O
  2. Input guardrailsScreen for prompt injection, PII and policy violations.
    I/O
  3. Plan and retrieveOrchestrator decides on retrieval; hybrid search + rerank return passages (L2, L3).
    I/O
  4. Assemble the promptSystem instructions, tools, retrieved context, history; cached prefix reused.
    I/O
  5. PrefillAll prompt tokens processed in parallel; KV cache written. Sets TTFT (L6).
    compute-bound
  6. DecodeOne token per step, streamed to the client. Sets TPOT (L6).
    memory-bound
  7. Tool calls (if any)Model emits a call, orchestrator runs it, result is appended, decode resumes.
    loop
  8. Output guardrailsCheck grounding, leakage and policy before release.
    I/O
  9. Log, trace and evaluateTrace, token counts, cost and feedback flow to observability and eval sets.
    async
Decisions architects make

Key trade-offs

DecisionOptionsChoose based on
KnowledgeRAG vs fine-tuningRAG for facts that change or need citations; fine-tuning for style, format and narrow skills.
Model sizeFrontier vs small / distilledTask difficulty, latency SLO and cost per request. Route easy traffic to small models.
HostingManaged API vs self-hosted open weightsData residency, volume (self-hosting pays off at high steady load), team capacity.
PrecisionFP16 vs FP8 / INT8 / INT4Memory and speed gains against measured quality loss on your eval set.
Control flowFixed chain vs autonomous agentPredictability and auditability vs flexibility. Start with chains, add agency where it earns its cost.
ContextLong context vs retrievalLong context is simpler but costs more per call and can dilute attention; retrieval scales to large corpora.
BatchingLatency vs throughputBigger batches raise tokens/s per GPU but increase per-request latency.