AI engineering glossary
What the jargon actually means — 250 terms from activation checkpointing to zero-shot, each with the engineering reality behind the buzzword.
- AI Risk Assessment
- Security & governance
- A documented analysis of how an AI system can affect people, organizations, and environments, including context, hazards, likelihood, impact, controls, residual risk, and monitoring responsibilities.
- Activation Checkpointing
- Math & training
- A training-memory technique that saves only selected forward-pass activations and recomputes the omitted ones during backpropagation.
- Activation Function
- Math & training
- A function applied after a linear or affine layer that introduces nonlinearity. Without it, composing layers with weights and biases collapses to one affine transformation. ReLU, GELU, and SiLU are common choices. The choice directly affects whether gradients flow during training.
- Adam (Optimizer)
- Math & training
- Adaptive Moment Estimation. It combines an exponential average of gradients with an exponential average of squared gradients, applies bias correction, and adapts the update scale per parameter. It is a useful baseline, but it still needs a suitable learning rate and schedule.
- AdamW
- Math & training
- An Adam variant that decouples weight decay from the gradient-based parameter update. That makes the shrinkage behavior easier to reason about than adding an L2 penalty inside Adam's adaptively scaled gradient.
- Admission Control
- Reliability & operations
- A pre-acceptance gate that decides whether a request may enter a bounded queue or service under the system's current capacity, priority, and policy.
- Agent
- Agents & tools
- A software system that lets a model select actions toward a goal, observe tool or environment results, and continue under an orchestration policy. An agent may use a loop, a state machine, a workflow engine, or human approvals. The model is one component, not the entire system.
- Agent Harness
- Agents & tools
- The runtime around a model that assembles context, exposes tools, manages state, enforces limits, records traces, and decides when the agent should continue, retry, ask, or stop.
- Agent Memory
- Agents & tools
- Information stored outside the model and selected for use in later agent steps, such as prior decisions, user preferences, task episodes, or verified facts.
- Agent Skill
- Agents & tools
- A discoverable directory of procedural instructions whose entry point is `SKILL.md`, with optional references, scripts, and assets that a compatible runtime can load in stages.
- Agent State
- Agents & tools
- The explicit data an agent carries across steps, such as the current objective, completed actions, tool results, open questions, budgets, approvals, and artifact references.
- Alignment
- Evaluation & safety
- The effort to make a model or AI system behave in ways that match intended goals, constraints, and human preferences across both expected and adversarial situations.
- Approval Gate
- Agents & tools
- A control point that blocks a consequential action until an authorized person or policy grants permission.
- Approximate Nearest Neighbor (ANN)
- Retrieval & generation
- A search method that returns vectors likely to be among the nearest to a query without exhaustively comparing the query with every stored vector.
- Attention
- Models & inference
- A mechanism that forms contextual representations by comparing query vectors with key vectors, normalizing the resulting scores, and using them to combine value vectors. Masks, position rules, or sparse patterns can restrict which positions participate.
- Audio Token
- Multimodal systems
- A discrete identifier produced by an audio codec or tokenizer for a short segment or feature of an audio signal, sometimes across several codebooks.
- Audit Log
- Security & governance
- A durable, access-controlled record of security- or accountability-relevant events, including who or what acted, what changed, when it happened, and the resulting status.
- Autograd
- Math & training
- A system that records or transforms tensor operations so it can compute derivatives, usually with reverse-mode automatic differentiation. You write the forward computation and the framework derives the gradients needed for backpropagation.
- Automatic Speech Recognition (ASR)
- Multimodal systems
- The task and system pipeline that maps a speech signal to a transcription, often with optional token or segment timing and confidence information.
- Autoregressive
- Models & inference
- A factorization in which each output token is predicted from the tokens that precede it. During generation, the selected token is appended to the sequence and becomes part of the next prediction's context.
- Autoscaling
- Infrastructure & serving
- A control loop that changes the number or capacity of serving workers from observed demand, resource use, or application metrics within configured bounds.
- Availability
- Reliability & operations
- The proportion of eligible service interactions or time windows in which users can obtain the defined acceptable service under a stated measurement boundary.
- BM25
- Retrieval & generation
- A lexical ranking function that scores a document from query-term matches while accounting for term rarity, repeated occurrences, and document length.
- Backpressure
- AI-native development
- A flow-control mechanism that slows or rejects upstream work when a downstream component cannot process it safely at the current rate.
- Backpropagation
- Math & training
- An efficient application of the chain rule that propagates derivatives from a scalar loss backward through a computation graph. It computes gradients; an optimizer uses those gradients to update parameters.
- Batch Size
- Math & training
- The number of examples whose losses contribute to one gradient estimate before an optimizer update. Larger batches can improve hardware utilization and reduce gradient noise, but they require more memory and may need different learning-rate or scheduling choices.
- Benchmark Contamination
- Evaluation & safety
- Overlap or information leakage between evaluation examples and data used to pretrain, tune, prompt, select, or otherwise improve the evaluated system.
- Byte Pair Encoding (BPE)
- Data & representations
- A subword-tokenization method that repeatedly merges frequent adjacent units to construct a fixed vocabulary from training text.
- CNN (Convolutional Neural Network)
- Models & inference
- A neural network that uses convolution operations (sliding filters over the input) to detect local patterns. Stacking convolutions detects increasingly complex features: edges, textures, objects.
- CUDA
- Models & inference
- NVIDIA's platform and programming model for general-purpose computation on compatible GPUs. Deep-learning frameworks use CUDA libraries and kernels to execute many tensor operations in parallel.
- Calibration
- Evaluation & safety
- The agreement between a system's stated confidence and the observed frequency with which predictions at that confidence are correct.
- Canary Release
- Reliability & operations
- A deployment strategy that exposes a new version to a limited slice of traffic or infrastructure before expanding the rollout.
- Chain of Thought (CoT)
- Prompting & context
- Intermediate reasoning used to decompose a task before producing an answer. A prompt can request a visible rationale, while some systems use internal reasoning that is not returned to the user.
- Checkpoint
- Agents & tools
- A durable snapshot used to resume from a known boundary. In a workflow, it stores operational state and artifact references. In model training, it can store parameters, optimizer state, scheduler state, and the training position.
- Chunked Prefill
- Infrastructure & serving
- A serving technique that divides a long prompt's prefill work into smaller schedulable pieces so prompt processing can interleave with decode work from other requests.
- Chunking
- Retrieval & generation
- Dividing source material into retrievable units before indexing. Chunk boundaries, overlap, metadata, and document structure determine whether retrieval returns enough context without flooding the prompt.
- Circuit Breaker
- AI-native development
- A reliability control that temporarily stops calls to a dependency after failures cross a threshold, then probes whether the dependency has recovered.
- Coding Agent
- AI-native development
- An agent specialized for software work that can inspect a repository, edit files, run development tools, and use their outputs to advance a scoped engineering task.
- Compensating Action
- Agents & tools
- A deliberate operation that semantically counteracts a completed side effect when the original operation cannot be rolled back atomically.
- Content Provenance
- Security & governance
- Verifiable information about the origin and editing history of a piece of media or other digital content, including the actors, tools, transformations, and assertions attached to it.
- Context Compression
- Prompting & context
- Reducing the token footprint of source material while attempting to preserve the information required for a later model decision.
- Context Engineering
- Prompting & context
- Designing the full information environment supplied to a model at each step, including instructions, selected files, retrieved evidence, tool results, examples, state, and output constraints.
- Context Window
- Prompting & context
- The maximum token capacity available to one model inference under a specific model and API contract. The capacity may include system instructions, messages, retrieved content, tool exchanges, and generated output, with provider-specific accounting and output limits.
- Continuous Batching
- Infrastructure & serving
- A serving scheduler that adds and removes generation requests at iteration boundaries instead of waiting for every request in a fixed batch to finish.
- Contrastive Learning
- Math & training
- Training by pulling similar pairs closer and pushing dissimilar pairs apart in embedding space. CLIP uses this: matching image-text pairs vs non-matching ones.
- Cosine Similarity
- Data & representations
- The normalized dot product of two vectors. It compares their direction rather than their magnitude and ranges from -1 to 1 for real-valued vectors.
- Cost per Successful Task
- AI-native development
- Total system cost divided by the number of tasks that satisfy a defined success criterion, including retries, failed runs, tool use, and evaluation overhead.
- Cross-Attention
- Multimodal systems
- Attention in which the query representation comes from one sequence or representation while keys and values come from another.
- Cross-Entropy
- Math & training
- A loss based on the negative log probability assigned to the target outcome. In next-token training, it penalizes the model when it assigns low probability to the observed next token.
- DPO (Direct Preference Optimization)
- Math & training
- A preference-optimization objective that trains a policy directly from preferred and rejected response pairs relative to a reference policy. It avoids running an explicit reward model and reinforcement-learning loop during this stage.
- Data Augmentation
- Math & training
- Creating modified examples, such as transformed images, perturbed audio, or paraphrased text, to increase training diversity without collecting entirely new source data. It can reduce overfitting when the transformation preserves the task signal.
- Data Classification
- Security & governance
- Assigning data to documented sensitivity or impact classes so handling, access, retention, sharing, and incident rules follow the consequences of disclosure or loss.
- Data Deduplication
- Data & representations
- Detecting and removing exact and near-duplicate examples within or across datasets.
- Data Exfiltration
- Security & governance
- Unauthorized transfer of protected data from a system or trust zone to a person, tool, service, or storage location that is not permitted to receive it.
- Data Leakage
- Data & representations
- Unintended use of information during training or feature construction that would not be available at the real prediction point or belongs to a held-out evaluation boundary.
- Data Lineage
- Security & governance
- A record of how a data artifact was derived across sources, transformations, joins, filters, versions, and downstream uses.
- Data Minimization
- Security & governance
- For personal data, limiting what is collected, processed, exposed, and retained to what is necessary for a specified purpose. Teams can apply the same discipline to sensitive non-personal data as an engineering control.
- Data Provenance
- Data & representations
- Traceable information about where data originated, who or what transformed it, which versions were used, and how derived artifacts relate to their sources.
- Dataset Split
- Data & representations
- A documented partition of examples into separate subsets for fitting, development decisions, and final evaluation.
- Datasheet for Datasets
- Security & governance
- Structured documentation of a dataset's motivation, composition, collection process, preprocessing, uses, distribution, maintenance, and known limitations.
- Deadline Propagation
- Reliability & operations
- Passing the remaining end-to-end time budget to downstream calls so each dependency knows how long the original request can still usefully wait.
- Decode Phase
- Infrastructure & serving
- The iterative stage of autoregressive inference that generates new tokens one step at a time after the input prefix has been processed.
- Decoder
- Models & inference
- A component that maps a representation into an output. In an encoder-decoder transformer, the decoder uses masked self-attention and cross-attention to generate outputs. Decoder-only language models instead generate from a single causal stack.
- Decoding Strategy
- Models & inference
- The algorithm that converts a model's sequence of next-token scores into selected tokens and a completed output.
- Defense in Depth
- Security & governance
- Using independent preventive, detective, and corrective controls at several system boundaries so one failed control does not determine the outcome.
- Delegation
- Agents & tools
- Assigning a bounded subtask to another person or agent together with the needed context, authority, output contract, and return conditions.
- Dense Retrieval
- Retrieval & generation
- First-stage retrieval that embeds queries and candidates into vector representations and ranks candidates by a similarity function.
- Diffusion Model
- Models & inference
- A generative model trained around a progressive noising process and a learned reverse process. Sampling usually begins from noise and applies repeated denoising steps, sometimes in a learned latent space.
- Disaggregated Serving
- Infrastructure & serving
- A serving architecture that runs prefill and decode work in separately provisioned worker pools and transfers the required attention state between them.
- Distribution Shift
- Evaluation & safety
- A difference between the data distribution used to build or evaluate a system and the distribution it encounters after deployment.
- Dropout
- Math & training
- During training, randomly setting a fraction of activations to zero encourages the network not to rely on one activation path. It is normally disabled for standard inference, although Monte Carlo dropout deliberately keeps it active to estimate uncertainty.
- Durable Execution
- Agents & tools
- Running a workflow so its state and completed steps survive process crashes, restarts, or long waits without redoing confirmed side effects.
- Dynamic Batching
- Infrastructure & serving
- A runtime policy that forms inference batches from queued requests according to compatible shapes, maximum size, priority, and allowed queue delay.
- Early Fusion
- Multimodal systems
- Combining raw or low-level representations from several modalities before most task-specific modeling occurs.
- Eigenvalue
- Math & training
- A scalar that describes how a linear transformation scales a corresponding nonzero eigenvector without changing its direction. In covariance-matrix PCA, larger eigenvalues correspond to directions with more variance.
- Embedding
- Data & representations
- A learned mapping from discrete items (words, images, users) to dense vectors in continuous space, where similar items end up close together
- Encoder
- Models & inference
- A component that transforms input into a representation. A transformer encoder commonly uses non-causal self-attention, subject to any masks, so each position can incorporate context from across the input.
- Epoch
- Math & training
- One traversal of the defined training dataset. In distributed or sampled training, the exact implementation of an epoch depends on the data loader and sampling policy.
- Error Budget
- Reliability & operations
- The amount of unsuccessful service allowed by a service-level objective over its measurement window before the objective is exhausted.
- Eval Set
- also: Evaluation set
- Evaluation & safety
- A versioned collection of inputs, expected properties, scoring rules, and metadata used to measure an AI system against a defined capability or risk.
- Evaluation (Eval)
- also: Eval
- Evaluation & safety
- A defined process for measuring model or system behavior on representative tasks using explicit success criteria, data, scorers, and review procedures.
- Exact Match (EM)
- Evaluation & safety
- A metric that counts an output as correct only when its normalized representation exactly equals an accepted reference answer.
- Expert Parallelism
- Infrastructure & serving
- Distributing mixture-of-experts subnetworks across devices and routing each token's activations to the devices that host its selected experts.
- Feature
- Data & representations
- An individual measurable property of the data. In classical ML, you engineer features by hand. In deep learning, the network learns features automatically from raw data.
- Few-Shot
- Prompting & context
- In-context learning that includes a small set of demonstrations before the target input so the model can infer the desired task, format, or decision boundary.
- Fine-tuning
- Math & training
- Continuing training from pretrained parameters on a narrower dataset or objective. Depending on the method, you may update all parameters, selected parameters, or added adapter parameters.
- Flaky Test
- AI-native development
- A test that can pass and fail across equivalent runs without a relevant change to the code or intended test environment.
- FlashAttention
- Infrastructure & serving
- An exact attention algorithm that tiles the computation to reduce transfers between accelerator memory levels while avoiding materialization of the full attention matrix in high-bandwidth memory.
- Function Calling
- Agents & tools
- A provider or application interface through which a model emits a structured request naming a tool and its arguments. Application code validates the request, performs the operation, and can return the result for another model step.
- GAN (Generative Adversarial Network)
- Models & inference
- A generator network tries to create realistic data while a discriminator network tries to tell real from fake. They train together: the generator gets better at fooling the discriminator, and the discriminator gets better at detecting fakes.
- GPT
- Models & inference
- Generative Pre-trained Transformer, a family label for generative transformer models pretrained on sequence-prediction objectives and adapted for downstream use. Product names and model architectures should not be treated as interchangeable.
- Goodput
- Infrastructure & serving
- The rate of completed requests that satisfy defined service constraints, such as both time-to-first-token and per-token latency objectives, under a stated workload.
- Graceful Degradation
- Reliability & operations
- Preserving a bounded core service when capacity or dependencies are impaired by reducing optional quality, features, freshness, or workload instead of failing every request.
- Gradient
- Math & training
- A vector of partial derivatives pointing in the direction of steepest increase. In ML, you go opposite to the gradient (gradient descent) to minimize the loss.
- Gradient Accumulation
- Math & training
- Summing or averaging gradients from several microbatches before performing one optimizer update.
- Gradient Clipping
- Math & training
- Limiting gradient values or their combined norm before an optimizer update when they exceed a chosen threshold.
- Gradient Descent
- Math & training
- A family of optimization updates that move parameters using the negative gradient of an objective, usually estimated from batches rather than the entire dataset.
- Grounding
- Retrieval & generation
- Connecting a generated answer or action to evidence, state, or observations that the system can identify and check.
- Guardrails
- Evaluation & safety
- System controls that constrain inputs, tool use, outputs, permissions, and escalation. They can include schemas, policy checks, classifiers, allowlists, sandboxing, approvals, and post-action verification.
- HNSW
- also: Hierarchical Navigable Small World
- Retrieval & generation
- An approximate-nearest-neighbor index that organizes vectors in layered proximity graphs and searches from coarse upper layers toward detailed lower layers.
- Hallucination
- Evaluation & safety
- Generated content that is false, unsupported by the available evidence, or inconsistent with the task's source of truth. It can arise even when the output is fluent and the model is not attempting to deceive.
- Handoff
- AI-native development
- A structured transfer of a task between people or agents that preserves the objective, current state, evidence, decisions, constraints, and remaining work.
- Human-in-the-Loop (HITL)
- also: Human oversight, human review
- Agents & tools
- A workflow design in which a person supplies judgment, correction, approval, or escalation at defined points in an AI-driven process.
- Hybrid Retrieval
- Retrieval & generation
- Retrieval that combines signals from different methods, commonly lexical matching and dense-vector similarity, before merging or reranking results.
- Hyperparameter
- Math & training
- A configuration choice that shapes model structure, optimization, data processing, or inference rather than being learned as an ordinary model parameter. Examples include learning rate, batch size, layer count, and decoding settings.
- Idempotency
- AI-native development
- The property that repeating the same operation with the same identity does not create additional side effects beyond the first successful application.
- Image Token
- Multimodal systems
- A model-specific visual unit represented as a vector or discrete code, commonly derived from an image patch, region, or learned visual-codebook entry.
- In-Context Learning
- Prompting & context
- A model adapting its behavior from instructions, examples, or patterns supplied in the current input without an ordinary parameter update.
- Incident Response
- Reliability & operations
- The coordinated process for detecting, analyzing, containing, recovering from, communicating, and learning from an event that threatens service, data, safety, or security.
- Indirect Prompt Injection
- Security & governance
- A prompt-injection attack delivered through content the system retrieves or observes, such as a webpage, document, email, image text, or tool result, rather than directly through the user's instruction.
- Inductive Bias
- Models & inference
- Structural or statistical assumptions that favor some functions or representations over others. Convolution favors locality and shared filters; causal masking favors prediction from preceding positions.
- Inference
- Models & inference
- Executing a trained model to produce predictions, scores, embeddings, or generated tokens without performing an ordinary training update to its parameters.
- Instruction Following
- Prompting & context
- A model capability to map natural-language directions and supplied context to behavior that satisfies the stated task and constraints.
- Instruction Hierarchy
- Prompting & context
- A rule set for resolving conflicts among instructions from sources with different authority, such as application policy, users, and untrusted retrieved content.
- Inter-Token Latency (ITL)
- Infrastructure & serving
- The elapsed time between two consecutive output-token arrival events for one request, calculated as `t_i - t_(i-1)` for an output token after the first.
- JAX
- Math & training
- A Python library for transforming numerical functions with automatic differentiation, compilation, vectorization, and parallel execution across accelerators. Its transformations work best with explicit state and functional-style code.
- Jailbreak
- Security & governance
- An adversarial input or interaction strategy intended to make a model produce behavior that its training or application controls are designed to prevent.
- KV Cache
- Models & inference
- Stored key and value tensors from earlier positions in autoregressive generation. Reusing them avoids recomputing attention projections for the unchanged prefix at every decoding step.
- Knowledge Distillation
- Math & training
- Training a student model to reproduce selected behavior or output distributions from a more capable teacher, often alongside ordinary target labels.
- LLM (Large Language Model)
- Models & inference
- A language model with enough capacity and broad training to perform many language tasks through prompting or adaptation. Most current LLMs use transformer architectures and sequence-prediction objectives, but size thresholds, data sources, and training recipes vary.
- LLM-as-a-Judge
- Evaluation & safety
- Using a language model to score, compare, classify, or critique another system's output against a rubric.
- Late Fusion
- Multimodal systems
- Processing modalities through separate encoders or predictors and combining their high-level representations, scores, or decisions near the task output.
- Latent Space
- Data & representations
- A learned representation space whose coordinates encode factors useful to a model. It may be lower-dimensional than the input, but compression is not required for every latent representation.
- Learning Rate
- Math & training
- A scale factor used by an optimizer to control parameter-update magnitude. Values that are too large can destabilize training; values that are too small can make useful progress impractically slow.
- Learning Rate Schedule
- Math & training
- A policy that changes the optimizer's learning rate as training progresses according to steps, epochs, metrics, or a predefined curve.
- Least Privilege
- Evaluation & safety
- Giving a model, agent, tool, or user only the permissions required for the current task, for only as long as those permissions are needed.
- LoRA (Low-Rank Adaptation)
- Math & training
- A method that keeps base weights frozen and learns low-rank update matrices for selected layers. It reduces the number of trainable parameters and can lower training memory relative to full-parameter fine-tuning.
- Load Shedding
- Reliability & operations
- Deliberately rejecting, dropping, or cancelling selected work at one or more overload boundaries when demand exceeds the capacity available to produce useful results.
- Logits
- Models & inference
- The model's unnormalized numeric scores for candidate outcomes before a normalization function or decoding rule converts them into selections.
- Loss Function
- Math & training
- An objective that maps predictions and targets, sometimes with regularization terms, to a value optimization tries to reduce. The loss determines which errors training directly rewards or penalizes.
- Lost in the Middle
- Prompting & context
- A long-context failure pattern in which model performance changes with evidence position and can degrade when relevant information sits between the beginning and end.
- MCP (Model Context Protocol)
- Agents & tools
- An open JSON-RPC protocol for a host to connect to servers that expose tools, resources, prompts, and extensions through defined request, result, discovery, and transport contracts. In revision 2026-07-28, every request carries its protocol version and client capabilities instead of relying on an initialization handshake or protocol session.
- Maximum Marginal Relevance (MMR)
- Retrieval & generation
- A selection rule that balances relevance to the query with novelty relative to items already selected.
- Membership Inference
- Security & governance
- An attack that estimates whether a particular record or example was included in a model's training data by observing model outputs or other accessible signals.
- Mixed Precision
- Math & training
- A numerical strategy that uses different data types for different operations, often lower precision for many matrix operations and higher precision for values that need more range or stability.
- MoE (Mixture of Experts)
- Models & inference
- An architecture with multiple expert subnetworks and a learned router that selects a subset for each input unit, often each token. Sparse activation can increase total parameter capacity without using every expert on every forward pass.
- Modality
- Multimodal systems
- A form of information with its own structure and acquisition process, such as text, image, audio, video, depth, or sensor measurements.
- Modality Alignment
- Multimodal systems
- Learning or establishing correspondences between representations from different modalities so semantically or temporally related items can be matched.
- Model Card
- Evaluation & safety
- A structured report describing a model's intended uses, evaluation conditions, performance characteristics, limitations, and relevant ethical or safety considerations.
- Model Router
- AI-native development
- A component that selects a model or provider for a request using requirements such as capability, latency, cost, context size, policy, and current availability.
- Model Serving
- Infrastructure & serving
- The runtime and API layer that loads versioned model artifacts, accepts inference requests, schedules execution, manages resources, and returns results under an operational contract.
- Multi Round-Trip Request (MRTR)
- also: MRTR
- Agents & tools
- An MCP request pattern in which an operation returns `resultType: input_required` with one or more `inputRequests`, then the client retries the original method with `inputResponses` and the exact returned `requestState`.
- Multimodal Fusion
- Multimodal systems
- Combining evidence or learned representations from more than one modality to produce a joint representation, prediction, or generated output.
- Multimodal Model
- Multimodal systems
- A model that learns from, relates, or generates more than one modality through representation, alignment, fusion, translation, or coordinated prediction.
- NaN (Not a Number)
- Math & training
- A floating-point value representing an undefined or unrepresentable numerical result. In training, NaNs can come from invalid operations, overflow, unstable normalization, excessive updates, or earlier corrupted values.
- Normalization
- Math & training
- A family of transformations that rescale or recenter inputs, activations, or features using defined statistics. Batch normalization and layer normalization use different axes and behave differently across training and inference.
- Nucleus Sampling (Top-p)
- also: Top-p sampling
- Models & inference
- A decoding method that samples from the smallest set of next-token candidates whose cumulative probability reaches a chosen threshold.
- Observability
- AI-native development
- The ability to understand an AI system's behavior from recorded inputs, outputs, state transitions, tool calls, timings, costs, errors, and evaluation signals.
- Optimizer
- Math & training
- An algorithm that transforms gradients into parameter updates. Plain stochastic gradient descent is a simple baseline; momentum, Adam, and other optimizers change the update using history or adaptive scaling. Each choice has different memory, stability, and tuning behavior.
- Orchestration
- Agents & tools
- The control logic that sequences, branches, delegates, retries, pauses, resumes, and terminates work across model and tool steps.
- Overfitting
- Math & training
- A generalization gap in which performance on training data is substantially better than performance on representative unseen data. Memorization can contribute, but the operational symptom is poor generalization.
- Paged KV Cache
- Infrastructure & serving
- A KV-cache memory manager that stores attention state in fixed-size blocks and maps logical sequence positions to physical blocks instead of requiring one contiguous allocation per sequence.
- Parameter
- Models & inference
- A value learned during training, commonly a weight, bias, embedding element, or normalization parameter. Parameter count is one measure of model capacity, but it does not directly determine quality, memory, or serving cost.
- Pass@k
- Evaluation & safety
- Across a task set, the fraction of tasks for which at least one of k sampled candidates passes a defined correctness test.
- Patch
- AI-native development
- A reviewable representation of changes to one or more files, usually expressed as additions and deletions against a known base revision.
- Patch Embedding
- Multimodal systems
- A learned projection that converts an image patch into a fixed-width vector used as one element of a transformer input sequence.
- Perplexity
- Models & inference
- The exponentiated average negative log-likelihood under a stated tokenization and logarithm convention. Lower values mean the model assigned higher probability to the evaluated sequence.
- Pipeline Parallelism
- Infrastructure & serving
- Partitioning sequential groups of model layers across devices and moving microbatches or requests through those stages as a pipeline.
- Planning
- Agents & tools
- Constructing, selecting, or revising a sequence of actions and dependencies intended to move from the current state to a goal.
- Postmortem
- Reliability & operations
- A durable incident record that explains impact, detection, response, contributing conditions, recovery, and owned follow-up actions without assigning blame as a substitute for analysis.
- Precision & Recall
- Evaluation & safety
- Precision asks how many flagged items were correct; recall asks how many relevant items were found. When you change the decision threshold for one fixed scoring model, improving recall often lowers precision and vice versa. A better model can improve both. F1 is their harmonic mean.
- Prefill
- also: Prefill Phase
- Infrastructure & serving
- The initial inference stage that processes all supplied input tokens to produce their representations and the attention state required for subsequent autoregressive generation.
- Prefix Caching
- Infrastructure & serving
- Reusing KV-cache blocks produced for an identical eligible token prefix across requests so the serving runtime can skip repeated prefix computation.
- Progressive Disclosure
- AI-native development
- Supplying a person or model with the minimum useful context first, then revealing deeper detail when the task or evidence requires it.
- Prompt Cache
- Prompting & context
- Reuse of provider-side or application-side computation for an identical or eligible prompt prefix so repeated inference avoids some preprocessing work.
- Prompt Engineering
- Prompting & context
- Designing model-facing instructions, examples, constraints, and output requirements to improve behavior on a defined task.
- Prompt Injection
- Evaluation & safety
- An attack or failure mode in which untrusted content influences a model to disregard intended instructions, expose data, misuse tools, or take actions outside the user's goal. The content can arrive directly from a user or indirectly through retrieved pages, files, messages, or tool output.
- Prompt Sensitivity
- Prompting & context
- Variation in model output or measured performance caused by changes to prompt wording, order, formatting, or examples that preserve the intended task.
- Provenance Attestation
- Security & governance
- Authenticated, machine-readable metadata that binds an artifact to claims about how, where, when, and from which inputs it was produced.
- Purpose Limitation
- Security & governance
- For personal data, collecting and using it only for specified, explicit purposes unless a new use has an appropriate compatible or authorized basis.
- QLoRA
- Math & training
- A parameter-efficient fine-tuning method that keeps a pretrained base model frozen in a low-bit quantized representation while training LoRA adapters with higher-precision computation where needed.
- Quantization
- Models & inference
- Representing weights, activations, or caches with lower-precision formats to reduce memory, bandwidth, or compute cost. Methods differ in calibration, granularity, data type, and whether conversion happens before, during, or after training.
- RAG (Retrieval-Augmented Generation)
- Retrieval & generation
- A system pattern that retrieves evidence relevant to a request and supplies selected content to a generative model before it answers or acts. Retrieval can use lexical, vector, structured, or hybrid methods.
- RLHF (Reinforcement Learning from Human Feedback)
- Math & training
- A family of pipelines that uses human feedback to learn a reward or preference signal and then optimizes a model policy against that signal. Implementations vary and need not all use the same reinforcement-learning algorithm.
- ROUGE
- Evaluation & safety
- A family of metrics that compares generated text with reference text using units such as n-gram overlap or longest common subsequence.
- Rate Limit
- AI-native development
- A policy that caps requests, tokens, concurrent work, or another resource within a defined time or capacity window.
- ReAct
- Agents & tools
- An agent pattern that interleaves task reasoning, a concrete action, and an observation returned by the environment before deciding the next step.
- ReLU
- Math & training
- Rectified Linear Unit, defined as `f(x) = max(0, x)`. It is inexpensive and has a non-saturating positive branch, though zero gradients on negative inputs can create inactive units.
- Readiness Probe
- Reliability & operations
- A diagnostic that tells the traffic-routing layer whether a service instance is currently able to accept requests.
- Recall@K
- Retrieval & generation
- For one query, Recall@K is `|relevant items intersecting the top k| / |relevant items|`. A dataset score aggregates those per-query values under a stated rule.
- Reciprocal Rank Fusion (RRF)
- Retrieval & generation
- A rank-fusion method that combines several result lists by summing contributions that decrease with each item's rank in each list.
- Red Teaming
- Security & governance
- A structured adversarial testing process in which authorized testers seek failures using documented objectives, threat assumptions, cases, and evidence.
- Regression Test
- AI-native development
- A repeatable check that protects behavior known to work, especially after code, prompt, model, retrieval, or tool changes.
- Repository Instructions
- AI-native development
- Version-controlled guidance that tells coding agents how a repository is organized, which commands and conventions apply, what boundaries to respect, and how to verify work.
- Repository Map
- AI-native development
- A compact, maintained description of a repository's important directories, ownership boundaries, entry points, build commands, tests, generated files, and local instructions.
- Reproducible Build
- AI-native development
- A build whose declared source, environment, and instructions can be independently rerun to produce bit-for-bit identical specified artifacts.
- Reranker
- Retrieval & generation
- A second-stage model or scoring function that reorders a small candidate set using a richer comparison between the query and each candidate.
- Retry Budget
- Reliability & operations
- A bound on retry traffic, usually expressed relative to original requests or over a time window, that prevents retries from consuming unbounded capacity.
- Retry with Backoff
- AI-native development
- Repeating a failed transient operation after progressively longer delays, usually with randomized jitter and a strict retry limit.
- Reviewer Agent
- AI-native development
- An agent assigned to inspect another agent's artifact or decision against explicit criteria and return findings or a verdict.
- Rollback
- Reliability & operations
- Restoring a previously known deployment or configuration when the current release violates operational, quality, or safety criteria.
- SFT (Supervised Fine-Tuning)
- Math & training
- Fine-tuning a pretrained model on paired inputs and desired responses so it learns the demonstrated behavior under the training distribution.
- Sandbox
- Agents & tools
- An isolated execution environment that restricts an agent's access to files, processes, network destinations, credentials, and host resources.
- Saturation
- Reliability & operations
- The degree to which a constrained resource or service has exhausted its capacity, including queued work that cannot begin promptly.
- Scope Contract
- AI-native development
- A concrete agreement that defines a task's goal, allowed and forbidden surfaces, expected artifacts, verification requirements, and stopping conditions.
- Self-Attention
- Models & inference
- Attention in which queries, keys, and values are derived from the same sequence representation. Scaled similarity scores are normalized and used to combine values, subject to causal, padding, local, or other masks.
- Semantic Cache
- AI-native development
- A cache that reuses a previous result when a new request is judged sufficiently similar under a chosen representation and threshold.
- Semantic Search
- Retrieval & generation
- Retrieval that represents a query and candidates in an embedding space and ranks candidates using a vector-similarity function.
- Separation of Duties
- Security & governance
- Dividing conflicting responsibilities or authority across independent roles so one principal cannot complete a high-risk action without another authorized decision.
- Service Level Indicator (SLI)
- Reliability & operations
- A quantitative measure of service behavior at a defined user-relevant boundary, such as successful request ratio or latency below a threshold.
- Service Level Objective (SLO)
- Reliability & operations
- A target range or threshold for a service-level indicator over a stated population and measurement window.
- Shadow Traffic
- Reliability & operations
- A copy of live request traffic sent to a candidate system for observation while the candidate response remains outside the primary user response path. Because the copied request still executes, its side effects must be isolated.
- Skill Bundle
- Agents & tools
- The complete installable skill directory, including `SKILL.md` and every reference, script, asset, fixture, or companion file required by the workflow.
- Skill Catalog
- Agents & tools
- The compact model-visible inventory of eligible skills, usually containing routing metadata such as name, description, and an internal source identifier rather than every skill body.
- Skill Discovery
- Agents & tools
- A runtime pipeline that searches configured roots, identifies candidate skill directories, validates their package contract, attaches scope and provenance, resolves collisions, and publishes eligible catalog entries.
- Skill Invocation
- Agents & tools
- The runtime-mediated process in which an eligible human, model, application, or other skill selects a skill and causes its instructions to enter the working context.
- Softmax
- Math & training
- A function defined by `softmax(x_i) = exp(x_i) / sum(exp(x_j))`, implemented with numerical stabilization. Its outputs are positive and sum to one, so they can parameterize a categorical distribution.
- Software Bill of Materials (SBOM)
- also: SBOM
- Security & governance
- A structured inventory of software components and relationships associated with a product or artifact, often including versions, suppliers, licenses, and identifiers.
- Speculative Decoding
- Models & inference
- An inference method in which a cheaper draft process proposes several tokens and the target model scores those draft positions in parallel. In exact sampling variants, an acceptance and correction rule preserves the target model's output distribution.
- Stateless MCP
- Agents & tools
- The MCP 2026-07-28 request model in which every request carries the protocol version and client capabilities in `params._meta`, while results carry an explicit `resultType`; no protocol state is keyed by an initialization handshake, connection, or `Mcp-Session-Id`.
- Stochastic Gradient Descent (SGD)
- also: SGD
- Math & training
- An optimizer family that updates parameters from a gradient estimated on a sampled example or minibatch rather than the complete training dataset.
- Stop Sequence
- Models & inference
- An application-specified token or text pattern that causes generation to stop when the decoding system encounters it.
- Streaming
- Models & inference
- Delivering incremental response events before the complete result is ready. A stream may contain token text, structured deltas, tool-call arguments, usage metadata, or status events depending on the API.
- Structured Output
- Agents & tools
- Model output constrained or validated against a machine-readable schema so application code can consume fields without parsing free-form prose.
- Swarm
- Agents & tools
- A loosely coordinated multi-agent pattern in which local agent decisions and message exchange produce system-level behavior. The term is used inconsistently, so the actual topology, state ownership, and termination rules must be specified.
- System Prompt
- Prompting & context
- A provider-defined instruction message or configuration supplied by the application to establish behavior and constraints within that provider's instruction hierarchy.
- Tail Latency
- Reliability & operations
- The latency experienced by the slowest portion of requests, commonly summarized with a high percentile under a stated workload and time window.
- Temperature
- Models & inference
- A decoding parameter that rescales logits before a probability distribution is formed. Higher positive values usually flatten the distribution; lower positive values sharpen it.
- Tensor
- Data & representations
- A typed array with a shape, data type, and device placement that frameworks use to represent inputs, parameters, activations, and gradients. Automatic-differentiation metadata is framework- and operation-dependent, not an inherent property of every tensor.
- Tensor Parallelism
- Infrastructure & serving
- Partitioning tensor operations within a model layer across devices, with collective communication combining partial results during the layer computation.
- Termination Condition
- Agents & tools
- An explicit rule that ends or pauses an agent run when it succeeds, fails, exhausts a budget, reaches a safe boundary, or requires escalation.
- Test Oracle
- AI-native development
- The mechanism, specification, reference, invariant, or human judgment used to decide whether observed program behavior is correct.
- Threat Model
- Security & governance
- A documented account of protected assets, trust boundaries, potential adversaries, assumed capabilities, attack paths, impacts, and planned controls.
- Time per Output Token (TPOT)
- Infrastructure & serving
- For one request with `N > 1` output tokens, the average post-first-token interval: `(t_N - t_1) / (N - 1)`. System distributions then aggregate those per-request averages.
- Time to First Token (TTFT)
- also: TTFT
- Models & inference
- The elapsed time from submitting a generation request until the client receives the first output token or content event under a defined measurement boundary.
- Token
- Data & representations
- An integer identifier produced by a model-specific tokenizer from text, bytes, images, audio, or another input representation. A token can be a whole word, part of a word, punctuation, whitespace, a byte sequence, or a special control symbol.
- Token Budget
- Prompting & context
- An explicit allocation of token capacity across instructions, evidence, history, tool results, reasoning or working space, and output.
- Tokenization
- Data & representations
- Converting an input representation into the ordered token identifiers a specific model or tokenizer accepts.
- Tokens per Second (TPS)
- also: TPS, output token throughput
- Infrastructure & serving
- A throughput measure reporting how many output tokens a serving system produces per unit time under a stated scope and workload.
- Tool Contract
- Agents & tools
- The complete agreement for a tool boundary: purpose, typed inputs, outputs, validation, permissions, side effects, errors, timeouts, idempotency, and evidence returned to the caller.
- Top-k Sampling
- Models & inference
- A decoding method that restricts the next-token distribution to the k highest-scoring candidates, renormalizes their probabilities, and samples from that set.
- Trace
- AI-native development
- A correlated record of one request or task across model calls, retrieval, tools, state transitions, retries, approvals, and evaluations.
- Transfer Learning
- Math & training
- Starting from representations or parameters learned on one data distribution or objective and adapting them for another. The transferable components and update strategy depend on architecture and task.
- Transformer
- Models & inference
- A neural-network architecture built from attention, position information, feed-forward sublayers, residual connections, and normalization. Encoder, decoder, and encoder-decoder variants use different masks and information flows.
- Trust Boundary
- Security & governance
- An interface where data, instructions, identity, or authority crosses between components or principals that operate under different trust assumptions.
- Underfitting
- Math & training
- A model or training setup has insufficient effective capacity, optimization, features, or training signal to capture useful patterns in the training data.
- VAE (Variational Autoencoder)
- Models & inference
- A latent-variable model trained with a reconstruction objective and a regularization term that keeps an approximate posterior close to a chosen prior. The reparameterization estimator allows gradients through stochastic latent sampling.
- Vector Database
- Retrieval & generation
- A storage and indexing system that supports nearest-neighbor queries over vector representations, often with metadata filtering, persistence, and approximate indexes.
- Verification Gate
- Evaluation & safety
- A control point that blocks progress until defined evidence satisfies a correctness or quality criterion.
- Vision Transformer (ViT)
- Multimodal systems
- A vision architecture that represents an image as a sequence of patch embeddings with position information and processes that sequence with transformer encoder blocks.
- Vision-Language Model (VLM)
- Multimodal systems
- A model that learns relationships between, or jointly processes, visual and language representations for tasks such as retrieval, description, question answering, or grounded generation.
- Visual Grounding
- Multimodal systems
- Connecting a language expression to spatial evidence in an image or video, such as a region, object, mask, or tracked entity.
- Vocabulary
- Data & representations
- The finite mapping between token identifiers and the units a tokenizer can emit, including ordinary, byte-level, and special control tokens.
- Warmup
- Math & training
- An initial training phase in which the learning rate rises from a smaller value toward the main schedule's target value.
- Weight
- Math & training
- A trainable coefficient in a model transformation. Weights are usually organized into tensors, and optimization adjusts them to reduce the training objective.
- Weight Decay
- Math & training
- An update rule that reduces selected parameter magnitudes over training, often by multiplying weights by a shrinkage factor separate from the gradient update.
- Worktree
- AI-native development
- In Git, a working directory attached to a repository and branch or commit, with shared object storage but its own checked-out files and index.
- Zero Trust
- Security & governance
- A security model that grants no implicit trust from network location or asset ownership and instead evaluates each access request against identity, device, resource, policy, and current context.
- Zero-Shot
- Prompting & context
- Performing a task from instructions or task framing without including task-specific demonstrations in the immediate input.