Core architecture and models
These are the terms that decide which model I can even serve, and on what hardware. Transformers. Neural networks that use self-attention to process words in relation to one another, forming the base of modern language models. Why it matters in production: self-attention is why cost and latency grow with context length rather than staying flat — every token attends to the tokens before it, so a longer prompt is not proportionally more expensive, it is disproportionately more expensive to prefill. Half the capacity arguments in Scalable vector database and this site’s token budgeting trace back to this one property. Multimodal models. Systems that process and integrate multiple data types at once, including text, images, audio, and video. Why it matters in production: a document with a chart is not text to a multimodal model, and the token math changes. Image and audio inputs bill as many more tokens than the same content transcribed, so my per-request estimate has to be measured on real inputs, not on the text layer. Small language models (SLMs). Compact, highly efficient models designed to run quickly on edge devices and local hardware with lower resource use. Why it matters in production: they are the answer to “we need this on-device” or “we cannot pay frontier prices for a classifier”. The trade is quality on long-horizon reasoning, so I use them where the task is narrow and the volume is high — routing, extraction, intent classification — not for the hard final answer. Reasoning models. Advanced models engineered to execute multi-step logic and internal deliberation before answering complex math, coding, or logic tasks. Why it matters in production: the deliberation is billed as output tokens. A reasoning model can turn a 1,500-token answer into tens of thousands of generated tokens, so a workload-level latency and cost budget must be re-planned before I switch a route to one. This is the hidden cost behind the workload-aware model routing described in Sample Resume. Mixture of experts (MoE). An architecture that routes specific tasks to specialized sub-networks within a larger model, improving speed and efficiency. Why it matters in production: only a fraction of parameters activate per token, which is how a very large model serves at a modest per-token price. It also means total parameter memory and inference compute are separate sizing problems — I can be bandwidth-bound on weights and still have spare FLOPs.Retrieval and context
Everything here exists because the model’s knowledge is frozen and mine is not. Retrieval-augmented generation (RAG). A method that lets an LLM pull fresh facts from external databases before answering, reducing the need to retrain the model. Why it matters in production: it converts model dependency into an infrastructure dependency — an index with freshness, permissions, and recall characteristics I now own. Full argument in Why Need RAG if Gemini Can Answer, mechanics in RAGs and Complete RAG tutorial. Embeddings and vector databases. Numerical vectors that capture semantic meaning, stored in specialized vector databases to enable fast similarity searches. Why it matters in production: dimensionality is a RAM bill. A 768-dimension float32 vector is768 × 4 = 3,072 bytes, so 100M of them is 307 GB before any index overhead. The arithmetic and the multipliers are in Data size estimation.
Chain-of-thought (CoT). A prompting technique that encourages the model to write out its step-by-step reasoning process to solve hard problems.
Why it matters in production: it makes intermediate logic inspectable, which helps debugging, and it inflates output tokens — the expensive direction of the bill. I also never show raw chain-of-thought to end users; it leaks reasoning that can contain policy contradictions or retrieved private text.
KV cache. A memory optimization technique used during inference to store previous key-value calculations and speed up response generation.
Why it matters in production: it is the actual GPU constraint behind long-context latency. At 10M tokens the cache becomes a memory bottleneck and generation crawls, as measured in Why Need RAG if Gemini Can Answer. Sharing it across users is the entire mechanism of Caching LLM Chats to Quickly Answer User Queries Without RAG.
Quantization. A process that shrinks model size and memory usage by lowering the precision of its numerical weights.
Why it matters in production: it buys capacity with accuracy I have to measure. Weights and vectors both shrink roughly linearly with bits per value — 4x less memory from fp32 to int8 — and the cost is a recall or quality loss that only shows up in evaluation, never in a smoke test.
Systems and orchestration
These terms describe the layer around the model, which is where most production incidents happen. Agentic AI. Autonomous AI systems designed to plan, use tools, and execute multi-step workflows toward a specific goal without constant human prodding. Why it matters in production: an agent turns a single model call into N calls with unbounded N. I need a step budget, a timeout, and a circuit breaker before I need a cleverer prompt; the loop and its failure modes are what What to Expect in FDE GenAI Roles expects a candidate to teach. Model Context Protocol (MCP). An open standard that simplifies how AI models securely connect to external tools, data sources, and software. Why it matters in production: it moves the integration surface from “one bespoke function schema per model vendor” to a shared tool contract, which is what lets a tool be reused across agents. The security consequence is real: every registered tool is a callable capability with its own blast radius, and it inherits the prompt-injection problem below. Guardrails. Safety and compliance filters placed around an LLM to block toxic language, personal data leaks, and off-topic or illegal responses. Why it matters in production: they sit on the request path, so they add latency and can reject valid traffic. Guardrails without a measured false-block rate silently degrade the product; I want both the block rate and the PII-recall rate tracked as SLOs, not vibes. LLM-as-a-judge. Using a powerful secondary AI model to evaluate, score, and test the outputs of another model automatically. Why it matters in production: it is the only way to regression-test generation quality at CI speed, since deterministic asserts cannot grade fluency or groundedness. Its failure modes are well documented — position bias toward the first candidate, verbosity bias, and self-preference for the judge’s own style — so I calibrate the judge against a small human-labelled set instead of trusting the score.How the terms compose into one bill
The useful part of a glossary is not the definitions, it is knowing which knobs interact. A single production request touches most of the list above:
Read that table as a cost model. If I want latency down, the levers are prefill length and call count, not “a faster GPU” — the same conclusion the cached-context arithmetic reaches in Caching LLM Chats to Quickly Answer User Queries Without RAG. If I want cost per query down, the lever is retrieved tokens per prompt, which is a chunking and top-k decision, not a pricing negotiation.
Three of these terms — RAG, agentic AI, and MCP — are the ones I am asked to explain from first principles most often in interviews. If I cannot explain why each exists before naming it, I do not understand it well enough to design with it, which is exactly the criterion the FDE teaching round applies.