Skip to main content
This is a snapshot glossary, and the date is the point: it reflects the vocabulary in circulation on 30 September, the same September 2026 window in which the model comparisons in Why Need RAG if Gemini Can Answer and the pricing in User usage at scale were collected. Term lists go stale faster than code. In six months some of these will be table stakes and others will be marketing noise, so I keep the list dated rather than pretending it is evergreen. I use two tests before I let a term into an architecture document. First, can I state what breaks if I ignore it? Second, does it change a number I have to size — latency, memory, cost per query, recall? The terms below pass both, which is why each entry ends with a production consequence rather than a restatement of the definition. If I am reading a term list from another source, I check it against Piyush Ranjan’s top Gen-AI terms, Faculty’s essential GenAI terminology guide, Vishal Chauhan’s must-know GenAI terms, TeamAI’s terms everyone should know, Level Up’s 12 essential GenAI terms, and Pankaj Shakya’s LLM terminology explainer. None of them agree on the boundary between “architecture” and “feature”, which is why I have grouped mine by what I have to do about each one.

Core architecture and models

These are the terms that decide which model I can even serve, and on what hardware. Transformers. Neural networks that use self-attention to process words in relation to one another, forming the base of modern language models. Why it matters in production: self-attention is why cost and latency grow with context length rather than staying flat — every token attends to the tokens before it, so a longer prompt is not proportionally more expensive, it is disproportionately more expensive to prefill. Half the capacity arguments in Scalable vector database and this site’s token budgeting trace back to this one property. Multimodal models. Systems that process and integrate multiple data types at once, including text, images, audio, and video. Why it matters in production: a document with a chart is not text to a multimodal model, and the token math changes. Image and audio inputs bill as many more tokens than the same content transcribed, so my per-request estimate has to be measured on real inputs, not on the text layer. Small language models (SLMs). Compact, highly efficient models designed to run quickly on edge devices and local hardware with lower resource use. Why it matters in production: they are the answer to “we need this on-device” or “we cannot pay frontier prices for a classifier”. The trade is quality on long-horizon reasoning, so I use them where the task is narrow and the volume is high — routing, extraction, intent classification — not for the hard final answer. Reasoning models. Advanced models engineered to execute multi-step logic and internal deliberation before answering complex math, coding, or logic tasks. Why it matters in production: the deliberation is billed as output tokens. A reasoning model can turn a 1,500-token answer into tens of thousands of generated tokens, so a workload-level latency and cost budget must be re-planned before I switch a route to one. This is the hidden cost behind the workload-aware model routing described in Sample Resume. Mixture of experts (MoE). An architecture that routes specific tasks to specialized sub-networks within a larger model, improving speed and efficiency. Why it matters in production: only a fraction of parameters activate per token, which is how a very large model serves at a modest per-token price. It also means total parameter memory and inference compute are separate sizing problems — I can be bandwidth-bound on weights and still have spare FLOPs.

Retrieval and context

Everything here exists because the model’s knowledge is frozen and mine is not. Retrieval-augmented generation (RAG). A method that lets an LLM pull fresh facts from external databases before answering, reducing the need to retrain the model. Why it matters in production: it converts model dependency into an infrastructure dependency — an index with freshness, permissions, and recall characteristics I now own. Full argument in Why Need RAG if Gemini Can Answer, mechanics in RAGs and Complete RAG tutorial. Embeddings and vector databases. Numerical vectors that capture semantic meaning, stored in specialized vector databases to enable fast similarity searches. Why it matters in production: dimensionality is a RAM bill. A 768-dimension float32 vector is 768 × 4 = 3,072 bytes, so 100M of them is 307 GB before any index overhead. The arithmetic and the multipliers are in Data size estimation. Chain-of-thought (CoT). A prompting technique that encourages the model to write out its step-by-step reasoning process to solve hard problems. Why it matters in production: it makes intermediate logic inspectable, which helps debugging, and it inflates output tokens — the expensive direction of the bill. I also never show raw chain-of-thought to end users; it leaks reasoning that can contain policy contradictions or retrieved private text. KV cache. A memory optimization technique used during inference to store previous key-value calculations and speed up response generation. Why it matters in production: it is the actual GPU constraint behind long-context latency. At 10M tokens the cache becomes a memory bottleneck and generation crawls, as measured in Why Need RAG if Gemini Can Answer. Sharing it across users is the entire mechanism of Caching LLM Chats to Quickly Answer User Queries Without RAG. Quantization. A process that shrinks model size and memory usage by lowering the precision of its numerical weights. Why it matters in production: it buys capacity with accuracy I have to measure. Weights and vectors both shrink roughly linearly with bits per value — 4x less memory from fp32 to int8 — and the cost is a recall or quality loss that only shows up in evaluation, never in a smoke test.

Systems and orchestration

These terms describe the layer around the model, which is where most production incidents happen. Agentic AI. Autonomous AI systems designed to plan, use tools, and execute multi-step workflows toward a specific goal without constant human prodding. Why it matters in production: an agent turns a single model call into N calls with unbounded N. I need a step budget, a timeout, and a circuit breaker before I need a cleverer prompt; the loop and its failure modes are what What to Expect in FDE GenAI Roles expects a candidate to teach. Model Context Protocol (MCP). An open standard that simplifies how AI models securely connect to external tools, data sources, and software. Why it matters in production: it moves the integration surface from “one bespoke function schema per model vendor” to a shared tool contract, which is what lets a tool be reused across agents. The security consequence is real: every registered tool is a callable capability with its own blast radius, and it inherits the prompt-injection problem below. Guardrails. Safety and compliance filters placed around an LLM to block toxic language, personal data leaks, and off-topic or illegal responses. Why it matters in production: they sit on the request path, so they add latency and can reject valid traffic. Guardrails without a measured false-block rate silently degrade the product; I want both the block rate and the PII-recall rate tracked as SLOs, not vibes. LLM-as-a-judge. Using a powerful secondary AI model to evaluate, score, and test the outputs of another model automatically. Why it matters in production: it is the only way to regression-test generation quality at CI speed, since deterministic asserts cannot grade fluency or groundedness. Its failure modes are well documented — position bias toward the first candidate, verbosity bias, and self-preference for the judge’s own style — so I calibrate the judge against a small human-labelled set instead of trusting the score.

How the terms compose into one bill

The useful part of a glossary is not the definitions, it is knowing which knobs interact. A single production request touches most of the list above: Read that table as a cost model. If I want latency down, the levers are prefill length and call count, not “a faster GPU” — the same conclusion the cached-context arithmetic reaches in Caching LLM Chats to Quickly Answer User Queries Without RAG. If I want cost per query down, the lever is retrieved tokens per prompt, which is a chunking and top-k decision, not a pricing negotiation.
Three of these terms — RAG, agentic AI, and MCP — are the ones I am asked to explain from first principles most often in interviews. If I cannot explain why each exists before naming it, I do not understand it well enough to design with it, which is exactly the criterion the FDE teaching round applies.
Where to go deeper: RAGs for retrieval architecture, Complete RAG tutorial for a runnable end-to-end pipeline, Scalable vector database for ANN and quantization at billions of vectors, Why Need RAG if Gemini Can Answer for the long-context-versus-retrieval argument, and What to Expect in FDE GenAI Roles for how this vocabulary is actually probed in a loop.