Daily note
Each day adds one short note on a core architecture idea. Notes appear here once the first one is posted.
Application & experience
Where people meet the system: chat surfaces, copilots inside existing tools, batch jobs, and APIs other services call. This layer owns the contract with the user: what goes in, what comes back, how long they wait, and what happens when the model is wrong.
Components
- Chat UI, embedded copilots, batch pipelines
- API gateway, auth, rate limiting, quotas
- Streaming transport (SSE / WebSocket)
- Feedback capture (thumbs, edits, escalations)
Design decisions
- Stream tokens or return a complete answer
- Synchronous chat vs async job with a callback
- Human-in-the-loop checkpoints for high-impact actions
- How citations and confidence are shown
Orchestration & agents
The control logic between the user and the model. It assembles the prompt, decides whether to retrieve, calls tools, loops until a goal is met, and manages what the model remembers. Most product behaviour lives here rather than in the weights.
Components
- Prompt templates and system instructions
- Tool / function calling with JSON schemas
- Planner–executor loops, routers, sub-agents
- Short-term (context) and long-term (store) memory
- Tool protocols such as MCP for connecting systems
Patterns
- Single call: prompt in, answer out
- Chain: fixed sequence of calls
- Router: classify, then send to a specialist model or prompt
- Agent loop: reason → act (tool) → observe → repeat, with a step budget
Knowledge & retrieval
Retrieval-augmented generation (RAG) grounds answers in your own documents without retraining. Quality is decided mostly by chunking and ranking, not by the model.
Ingest path
- Parse (PDF, HTML, tables) and clean
- Chunk: typically 200–800 tokens with overlap, or by document structure
- Embed each chunk into a dense vector
- Index in a vector store (HNSW, IVF-PQ) with metadata filters
Query path
- Rewrite or expand the query
- Hybrid search: dense vectors + BM25 keywords
- Rerank top-k with a cross-encoder
- Pack the best passages into the prompt with source IDs
Foundation model
Almost every current large language model is a decoder-only transformer: a stack of identical blocks that turns a sequence of tokens into a probability distribution over the next token. Variants change the attention pattern, the feed-forward layer, or add encoders for images and audio.
Building blocks
- Tokenizer (BPE / SentencePiece), vocab ~32k–256k
- Token embeddings + rotary positions (RoPE)
- Self-attention: multi-head, grouped-query (GQA)
- Feed-forward (SwiGLU) or Mixture-of-Experts
- RMSNorm and residual connections
Variants to know
- Dense: every parameter used per token
- MoE: router picks top-k experts; large total, small active params
- Multimodal: vision/audio encoder projected into token space
- Long context: RoPE scaling, sliding-window or sparse attention
One transformer block
Each block reads the running representation (the residual stream), lets every token look back at earlier tokens through attention, then transforms each position independently with the feed-forward layer. Residual connections (dashed) add each sub-layer's output back to its input, which is what lets very deep stacks train.
- Attention mixes information across positions; cost grows with sequence length squared in prefill.
- Feed-forward holds most parameters (about two-thirds in a dense model) and much of the stored knowledge.
- GQA shares key/value heads across query heads to shrink the KV cache.
- MoE swaps the single feed-forward for many experts; only top-k run per token.
Training & adaptation
A model passes through stages, each with different data and objectives. For most enterprise work the choice is not whether to pretrain but which adaptation to apply on top of an existing model.
Stages
- Pretraining: next-token prediction on trillions of tokens
- Supervised fine-tuning: instruction–response pairs
- Preference tuning: RLHF (reward model + PPO) or DPO
- Distillation: a small model learns from a larger one
Adaptation ladder (cheapest first)
- Prompting and few-shot examples
- RAG for knowledge that changes
- Parameter-efficient tuning: LoRA / QLoRA adapters
- Full fine-tune for behaviour and format at scale
Inference & serving
Serving splits into two phases with different bottlenecks. Prefill processes the whole prompt in parallel and is compute-bound; it sets time-to-first-token. Decode produces one token at a time, rereads all weights and the KV cache each step, and is memory-bandwidth-bound; it sets time-per-output-token.
Techniques
- KV cache, with paged allocation (PagedAttention)
- Continuous (in-flight) batching
- Quantization: FP8, INT8, INT4 weights
- Speculative decoding with a draft model
- Prefix / prompt caching for shared system prompts
Metrics
- TTFT: time to first token (ms)
- TPOT: time per output token (ms)
- Throughput: tokens/s per GPU
- Goodput: requests/s that meet the latency SLO
- Cost per 1M input / output tokens
KV cache per token = 2 × layers × kv_heads × head_dim × bytes
e.g. 2 × 32 × 8 × 128 × 2 B = 128 KiB/token → 8,192-token context ≈ 1 GiB per sequence
Inference compute ≈ 2 · N FLOPs per generated token
The KV cache, explained
A language model writes one token at a time. To pick each new token, attention compares it with every earlier token using their keys (what each token offers to be matched on) and values (what it contributes once matched). Those keys and values never change once computed, so serving engines keep them in GPU memory instead of recomputing them. That store is the KV cache. It makes generation fast, and it is usually what limits how many people one GPU can serve.
1 · What the cache saves
To produce token 5, the model needs keys and values for tokens 1 to 5. Without a cache it recomputes all of them at every step. With a cache it computes only the newest one and reads the rest. Over a 1,000-token answer that is 500,500 key/value computations per layer without a cache, against 1,000 with one.
2 · What one token costs in memory
Every token in every active conversation holds this much, in every layer, for as long as the conversation is live. The shape shown is typical of an 8B-parameter model.
3 · How many conversations fit on one 80 GB GPU
Same GPU, same model. Only the context length changes, and capacity falls from 238 conversations to 3. Long context is a memory decision before it is a model decision.
Assumes an 8B model in FP16 (32 layers, 8 KV heads, head dim 128) and every conversation at full length. Ignores activations and runtime overhead, so real capacity is somewhat lower.4 · How architects shrink it
Make each token smaller
- Grouped-query attention: several query heads share one KV head
- FP8 / INT8 KV cache: halve the bytes per value
- Latent attention (MLA): store a compressed latent and expand on use
Store fewer tokens, waste less
- PagedAttention: allocate in small blocks, so little memory sits reserved but unused
- Prefix caching: a shared system prompt is stored once for all users
- Sliding-window attention: keep only the most recent N tokens in some layers
- Offload: move idle conversations to CPU memory or SSD
Data foundation
Every layer above depends on data that is clean, permissioned and traceable. The same pipelines feed pretraining corpora, fine-tuning sets, retrieval indexes and evaluation sets.
Components
- Ingestion from lakes, warehouses, SaaS sources
- Deduplication (exact + MinHash near-dup)
- Quality and toxicity filtering, PII redaction
- Labelling and synthetic data generation
Controls
- Lineage: which data trained or grounded which output
- Access control carried into the retrieval index
- Versioned datasets and eval sets
- Retention and consent rules
Compute & infrastructure
Accelerators, memory and the network between them. For large models, memory capacity and bandwidth matter as much as raw FLOPs, and the interconnect decides how well work splits across devices.
Hardware
- GPUs / TPUs / custom accelerators
- HBM capacity and bandwidth per device
- Intra-node links (NVLink) and inter-node fabric (InfiniBand, RoCE)
- Managed platforms: Azure AI Foundry, Amazon Bedrock / SageMaker, Vertex AI
Parallelism
- Data: copies of the model, split the batch
- Tensor: split each matrix across devices
- Pipeline: split layers into stages
- Expert: spread MoE experts across devices
- Sharded optimizers (ZeRO / FSDP)
Cross-cutting concerns
Evaluation
- Offline golden sets per use case
- LLM-as-judge with human calibration
- RAG: retrieval recall, faithfulness, answer relevance
- Online A/B tests and regression gates in CI
Safety & guardrails
- Input checks: prompt injection, PII, policy
- Output checks: grounding, toxicity, data leakage
- Tool permissions and least privilege for agents
- Red-teaming before each release
Observability
- Traces per request across retrieval, tools, model
- Token usage and cost attribution
- Latency (TTFT, TPOT) and error rates
- Drift in inputs and quality scores
Governance
- Model and prompt registry with versions
- Risk tiering and approvals per use case
- Audit logs and explainability records
- Regulatory mapping (e.g. EU AI Act, NIST AI RMF)
Identity & policy: on whose authority?
Compute and data platforms settled access control years ago with IAM: every call carries a principal, and policy decides what it may touch. AI systems reopen the question, because the caller is increasingly an agent acting for a person. Every layer, from retrieval to tool execution, now has to answer one question: on whose authority is this running, and with what limits?
1 · Authority should narrow at every hop
The agent never holds more access than the person it acts for, and each step down gets only what that step needs. If a token leaks or a prompt injection hijacks the agent, the damage stays inside the narrowest scope.
2 · The question each layer must answer
| Layer | Question | Control |
|---|---|---|
| L1 Application | Who is the user? | SSO with OIDC, MFA, session lifetime |
| L2 Orchestration | What may this agent do for them? | Delegated, short-lived tokens (OAuth on-behalf-of, token exchange), step budgets, human approval for high-impact actions |
| L2 Tools | What can this call change? | Per-tool scopes and allow-lists, OAuth on MCP servers, read-only by default, dry runs |
| L3 Retrieval | Which documents may this user see? | Source ACLs carried into the index and filtered at query time, before anything reaches the prompt |
| L6 Gateway | Which tenant, quota and model? | API gateway, tenant isolation, rate limits, data-residency routing |
| L7 Data | May this data be used for this purpose? | Purpose-based access, consent flags, lineage |
| Audit | Who did what, for whom? | Log the full chain per action: user → agent → tool → resource |
3 · Design rules
Governance by design: proving the system does what it should
Governance added at review time produces documents about a system. Governance built in at design time produces evidence of how it behaves. The rule is simple: every intended behaviour becomes a testable requirement, and every requirement leaves a trail that a reviewer or auditor can follow from intent to production.
1 · The traceability chain
- Intended useWhat the system is for, who uses it, what it must never do, and its risk tier.
- RequirementEach "must" and "must not" written so it can be measured.
- Design controlThe mechanism that enforces it: retrieval scope, guardrail, human approval, tool permission.
- Eval with a thresholdA test set and pass mark, run in CI. A failed gate blocks the release.
- Production monitorThe same check sampled on live traffic, with an alert and an owner.
- EvidenceVersioned eval reports, approvals, logs and incidents, produced by the pipeline itself.
Worked example, an HR policy assistant. Intent: answer employee questions from the approved handbook only. Requirement: no answer without a citation to the handbook. Design control: retrieval limited to the handbook index, and the response is rejected if it has no citation. Eval: 300 test questions, at least 95% grounded and zero uncited answers. Monitor: 2% of live answers scored for groundedness each week. Evidence: eval report per release, weekly monitor results, sign-off record.
2 · The evidence pack
| Evidence | Produced by | When | What frameworks ask for |
|---|---|---|---|
| Intended use & risk tier | Product owner with risk | Design | NIST AI RMF Map; EU AI Act risk classification |
| Requirements as eval sets | Engineering | Design and build | NIST AI RMF Measure; accuracy and robustness obligations |
| Model, data & prompt cards | Engineering | Every release | Technical documentation; data governance |
| Eval reports & gate decisions | CI pipeline | Every release | ISO/IEC 42001 performance evaluation |
| Human-oversight design & approvals | Accountable owner | Before deploy | Human oversight obligations |
| Logs, monitors & incidents | Platform team | Continuously | Record-keeping; NIST AI RMF Manage |
3 · The AI engineering governance skill set
Technical
- Eval design and metrics
- RAG, agents and guardrails
- CI/CD and MLOps pipelines
- Observability and logging
AI engineering governance
- Turn policy into testable requirements
- Set release gates and thresholds
- Design evidence that is generated, not assembled
- Explain model behaviour to auditors and the board
Governance
- Risk assessment and tiering
- Regulatory mapping
- Accountability and approvals
- Audit and assurance
The middle column is the scarce role: someone who can read a model card and a control framework with equal confidence.
Lifecycle of one request
- Receive and authenticateGateway checks identity, quota and tenant (L1).I/O
- Input guardrailsScreen for prompt injection, PII and policy violations.I/O
- Plan and retrieveOrchestrator decides on retrieval; hybrid search + rerank return passages (L2, L3).I/O
- Assemble the promptSystem instructions, tools, retrieved context, history; cached prefix reused.I/O
- PrefillAll prompt tokens processed in parallel; KV cache written. Sets TTFT (L6).compute-bound
- DecodeOne token per step, streamed to the client. Sets TPOT (L6).memory-bound
- Tool calls (if any)Model emits a call, orchestrator runs it, result is appended, decode resumes.loop
- Output guardrailsCheck grounding, leakage and policy before release.I/O
- Log, trace and evaluateTrace, token counts, cost and feedback flow to observability and eval sets.async
Key trade-offs
| Decision | Options | Choose based on |
|---|---|---|
| Knowledge | RAG vs fine-tuning | RAG for facts that change or need citations; fine-tuning for style, format and narrow skills. |
| Model size | Frontier vs small / distilled | Task difficulty, latency SLO and cost per request. Route easy traffic to small models. |
| Hosting | Managed API vs self-hosted open weights | Data residency, volume (self-hosting pays off at high steady load), team capacity. |
| Precision | FP16 vs FP8 / INT8 / INT4 | Memory and speed gains against measured quality loss on your eval set. |
| Control flow | Fixed chain vs autonomous agent | Predictability and auditability vs flexibility. Start with chains, add agency where it earns its cost. |
| Context | Long context vs retrieval | Long context is simpler but costs more per call and can dilute attention; retrieval scales to large corpora. |
| Batching | Latency vs throughput | Bigger batches raise tokens/s per GPU but increase per-request latency. |