AI engineering glossary

What the jargon actually means — 250 terms from activation checkpointing to zero-shot, each with the engineering reality behind the buzzword.

AI Risk Assessment
Security & governance
A documented analysis of how an AI system can affect people, organizations, and environments, including context, hazards, likelihood, impact, controls, residual risk, and monitoring responsibilities.
Activation Checkpointing
Math & training
A training-memory technique that saves only selected forward-pass activations and recomputes the omitted ones during backpropagation.
Activation Function
Math & training
A function applied after a linear or affine layer that introduces nonlinearity. Without it, composing layers with weights and biases collapses to one affine transformation. ReLU, GELU, and SiLU are common choices. The choice directly affects whether gradients flow during training.
Adam (Optimizer)
Math & training
Adaptive Moment Estimation. It combines an exponential average of gradients with an exponential average of squared gradients, applies bias correction, and adapts the update scale per parameter. It is a useful baseline, but it still needs a suitable learning rate and schedule.
AdamW
Math & training
An Adam variant that decouples weight decay from the gradient-based parameter update. That makes the shrinkage behavior easier to reason about than adding an L2 penalty inside Adam's adaptively scaled gradient.
Admission Control
Reliability & operations
A pre-acceptance gate that decides whether a request may enter a bounded queue or service under the system's current capacity, priority, and policy.
Agent
Agents & tools
A software system that lets a model select actions toward a goal, observe tool or environment results, and continue under an orchestration policy. An agent may use a loop, a state machine, a workflow engine, or human approvals. The model is one component, not the entire system.
Agent Harness
Agents & tools
The runtime around a model that assembles context, exposes tools, manages state, enforces limits, records traces, and decides when the agent should continue, retry, ask, or stop.
Agent Memory
Agents & tools
Information stored outside the model and selected for use in later agent steps, such as prior decisions, user preferences, task episodes, or verified facts.
Agent Skill
Agents & tools
A discoverable directory of procedural instructions whose entry point is `SKILL.md`, with optional references, scripts, and assets that a compatible runtime can load in stages.
Agent State
Agents & tools
The explicit data an agent carries across steps, such as the current objective, completed actions, tool results, open questions, budgets, approvals, and artifact references.
Alignment
Evaluation & safety
The effort to make a model or AI system behave in ways that match intended goals, constraints, and human preferences across both expected and adversarial situations.
Approval Gate
Agents & tools
A control point that blocks a consequential action until an authorized person or policy grants permission.
Approximate Nearest Neighbor (ANN)
Retrieval & generation
A search method that returns vectors likely to be among the nearest to a query without exhaustively comparing the query with every stored vector.
Attention
Models & inference
A mechanism that forms contextual representations by comparing query vectors with key vectors, normalizing the resulting scores, and using them to combine value vectors. Masks, position rules, or sparse patterns can restrict which positions participate.
Audio Token
Multimodal systems
A discrete identifier produced by an audio codec or tokenizer for a short segment or feature of an audio signal, sometimes across several codebooks.
Audit Log
Security & governance
A durable, access-controlled record of security- or accountability-relevant events, including who or what acted, what changed, when it happened, and the resulting status.
Autograd
Math & training
A system that records or transforms tensor operations so it can compute derivatives, usually with reverse-mode automatic differentiation. You write the forward computation and the framework derives the gradients needed for backpropagation.
Automatic Speech Recognition (ASR)
Multimodal systems
The task and system pipeline that maps a speech signal to a transcription, often with optional token or segment timing and confidence information.
Autoregressive
Models & inference
A factorization in which each output token is predicted from the tokens that precede it. During generation, the selected token is appended to the sequence and becomes part of the next prediction's context.
Autoscaling
Infrastructure & serving
A control loop that changes the number or capacity of serving workers from observed demand, resource use, or application metrics within configured bounds.
Availability
Reliability & operations
The proportion of eligible service interactions or time windows in which users can obtain the defined acceptable service under a stated measurement boundary.
BM25
Retrieval & generation
A lexical ranking function that scores a document from query-term matches while accounting for term rarity, repeated occurrences, and document length.
Backpressure
AI-native development
A flow-control mechanism that slows or rejects upstream work when a downstream component cannot process it safely at the current rate.
Backpropagation
Math & training
An efficient application of the chain rule that propagates derivatives from a scalar loss backward through a computation graph. It computes gradients; an optimizer uses those gradients to update parameters.
Batch Size
Math & training
The number of examples whose losses contribute to one gradient estimate before an optimizer update. Larger batches can improve hardware utilization and reduce gradient noise, but they require more memory and may need different learning-rate or scheduling choices.
Benchmark Contamination
Evaluation & safety
Overlap or information leakage between evaluation examples and data used to pretrain, tune, prompt, select, or otherwise improve the evaluated system.
Byte Pair Encoding (BPE)
Data & representations
A subword-tokenization method that repeatedly merges frequent adjacent units to construct a fixed vocabulary from training text.
CNN (Convolutional Neural Network)
Models & inference
A neural network that uses convolution operations (sliding filters over the input) to detect local patterns. Stacking convolutions detects increasingly complex features: edges, textures, objects.
CUDA
Models & inference
NVIDIA's platform and programming model for general-purpose computation on compatible GPUs. Deep-learning frameworks use CUDA libraries and kernels to execute many tensor operations in parallel.
Calibration
Evaluation & safety
The agreement between a system's stated confidence and the observed frequency with which predictions at that confidence are correct.
Canary Release
Reliability & operations
A deployment strategy that exposes a new version to a limited slice of traffic or infrastructure before expanding the rollout.
Chain of Thought (CoT)
Prompting & context
Intermediate reasoning used to decompose a task before producing an answer. A prompt can request a visible rationale, while some systems use internal reasoning that is not returned to the user.
Checkpoint
Agents & tools
A durable snapshot used to resume from a known boundary. In a workflow, it stores operational state and artifact references. In model training, it can store parameters, optimizer state, scheduler state, and the training position.
Chunked Prefill
Infrastructure & serving
A serving technique that divides a long prompt's prefill work into smaller schedulable pieces so prompt processing can interleave with decode work from other requests.
Chunking
Retrieval & generation
Dividing source material into retrievable units before indexing. Chunk boundaries, overlap, metadata, and document structure determine whether retrieval returns enough context without flooding the prompt.
Circuit Breaker
AI-native development
A reliability control that temporarily stops calls to a dependency after failures cross a threshold, then probes whether the dependency has recovered.
Coding Agent
AI-native development
An agent specialized for software work that can inspect a repository, edit files, run development tools, and use their outputs to advance a scoped engineering task.
Compensating Action
Agents & tools
A deliberate operation that semantically counteracts a completed side effect when the original operation cannot be rolled back atomically.
Content Provenance
Security & governance
Verifiable information about the origin and editing history of a piece of media or other digital content, including the actors, tools, transformations, and assertions attached to it.
Context Compression
Prompting & context
Reducing the token footprint of source material while attempting to preserve the information required for a later model decision.
Context Engineering
Prompting & context
Designing the full information environment supplied to a model at each step, including instructions, selected files, retrieved evidence, tool results, examples, state, and output constraints.
Context Window
Prompting & context
The maximum token capacity available to one model inference under a specific model and API contract. The capacity may include system instructions, messages, retrieved content, tool exchanges, and generated output, with provider-specific accounting and output limits.
Continuous Batching
Infrastructure & serving
A serving scheduler that adds and removes generation requests at iteration boundaries instead of waiting for every request in a fixed batch to finish.
Contrastive Learning
Math & training
Training by pulling similar pairs closer and pushing dissimilar pairs apart in embedding space. CLIP uses this: matching image-text pairs vs non-matching ones.
Cosine Similarity
Data & representations
The normalized dot product of two vectors. It compares their direction rather than their magnitude and ranges from -1 to 1 for real-valued vectors.
Cost per Successful Task
AI-native development
Total system cost divided by the number of tasks that satisfy a defined success criterion, including retries, failed runs, tool use, and evaluation overhead.
Cross-Attention
Multimodal systems
Attention in which the query representation comes from one sequence or representation while keys and values come from another.
Cross-Entropy
Math & training
A loss based on the negative log probability assigned to the target outcome. In next-token training, it penalizes the model when it assigns low probability to the observed next token.
DPO (Direct Preference Optimization)
Math & training
A preference-optimization objective that trains a policy directly from preferred and rejected response pairs relative to a reference policy. It avoids running an explicit reward model and reinforcement-learning loop during this stage.
Data Augmentation
Math & training
Creating modified examples, such as transformed images, perturbed audio, or paraphrased text, to increase training diversity without collecting entirely new source data. It can reduce overfitting when the transformation preserves the task signal.
Data Classification
Security & governance
Assigning data to documented sensitivity or impact classes so handling, access, retention, sharing, and incident rules follow the consequences of disclosure or loss.
Data Deduplication
Data & representations
Detecting and removing exact and near-duplicate examples within or across datasets.
Data Exfiltration
Security & governance
Unauthorized transfer of protected data from a system or trust zone to a person, tool, service, or storage location that is not permitted to receive it.
Data Leakage
Data & representations
Unintended use of information during training or feature construction that would not be available at the real prediction point or belongs to a held-out evaluation boundary.
Data Lineage
Security & governance
A record of how a data artifact was derived across sources, transformations, joins, filters, versions, and downstream uses.
Data Minimization
Security & governance
For personal data, limiting what is collected, processed, exposed, and retained to what is necessary for a specified purpose. Teams can apply the same discipline to sensitive non-personal data as an engineering control.
Data Provenance
Data & representations
Traceable information about where data originated, who or what transformed it, which versions were used, and how derived artifacts relate to their sources.
Dataset Split
Data & representations
A documented partition of examples into separate subsets for fitting, development decisions, and final evaluation.
Datasheet for Datasets
Security & governance
Structured documentation of a dataset's motivation, composition, collection process, preprocessing, uses, distribution, maintenance, and known limitations.
Deadline Propagation
Reliability & operations
Passing the remaining end-to-end time budget to downstream calls so each dependency knows how long the original request can still usefully wait.
Decode Phase
Infrastructure & serving
The iterative stage of autoregressive inference that generates new tokens one step at a time after the input prefix has been processed.
Decoder
Models & inference
A component that maps a representation into an output. In an encoder-decoder transformer, the decoder uses masked self-attention and cross-attention to generate outputs. Decoder-only language models instead generate from a single causal stack.
Decoding Strategy
Models & inference
The algorithm that converts a model's sequence of next-token scores into selected tokens and a completed output.
Defense in Depth
Security & governance
Using independent preventive, detective, and corrective controls at several system boundaries so one failed control does not determine the outcome.
Delegation
Agents & tools
Assigning a bounded subtask to another person or agent together with the needed context, authority, output contract, and return conditions.
Dense Retrieval
Retrieval & generation
First-stage retrieval that embeds queries and candidates into vector representations and ranks candidates by a similarity function.
Diffusion Model
Models & inference
A generative model trained around a progressive noising process and a learned reverse process. Sampling usually begins from noise and applies repeated denoising steps, sometimes in a learned latent space.
Disaggregated Serving
Infrastructure & serving
A serving architecture that runs prefill and decode work in separately provisioned worker pools and transfers the required attention state between them.
Distribution Shift
Evaluation & safety
A difference between the data distribution used to build or evaluate a system and the distribution it encounters after deployment.
Dropout
Math & training
During training, randomly setting a fraction of activations to zero encourages the network not to rely on one activation path. It is normally disabled for standard inference, although Monte Carlo dropout deliberately keeps it active to estimate uncertainty.
Durable Execution
Agents & tools
Running a workflow so its state and completed steps survive process crashes, restarts, or long waits without redoing confirmed side effects.
Dynamic Batching
Infrastructure & serving
A runtime policy that forms inference batches from queued requests according to compatible shapes, maximum size, priority, and allowed queue delay.
Early Fusion
Multimodal systems
Combining raw or low-level representations from several modalities before most task-specific modeling occurs.
Eigenvalue
Math & training
A scalar that describes how a linear transformation scales a corresponding nonzero eigenvector without changing its direction. In covariance-matrix PCA, larger eigenvalues correspond to directions with more variance.
Embedding
Data & representations
A learned mapping from discrete items (words, images, users) to dense vectors in continuous space, where similar items end up close together
Encoder
Models & inference
A component that transforms input into a representation. A transformer encoder commonly uses non-causal self-attention, subject to any masks, so each position can incorporate context from across the input.
Epoch
Math & training
One traversal of the defined training dataset. In distributed or sampled training, the exact implementation of an epoch depends on the data loader and sampling policy.
Error Budget
Reliability & operations
The amount of unsuccessful service allowed by a service-level objective over its measurement window before the objective is exhausted.
Eval Set
also: Evaluation set
Evaluation & safety
A versioned collection of inputs, expected properties, scoring rules, and metadata used to measure an AI system against a defined capability or risk.
Evaluation (Eval)
also: Eval
Evaluation & safety
A defined process for measuring model or system behavior on representative tasks using explicit success criteria, data, scorers, and review procedures.
Exact Match (EM)
Evaluation & safety
A metric that counts an output as correct only when its normalized representation exactly equals an accepted reference answer.
Expert Parallelism
Infrastructure & serving
Distributing mixture-of-experts subnetworks across devices and routing each token's activations to the devices that host its selected experts.
Feature
Data & representations
An individual measurable property of the data. In classical ML, you engineer features by hand. In deep learning, the network learns features automatically from raw data.
Few-Shot
Prompting & context
In-context learning that includes a small set of demonstrations before the target input so the model can infer the desired task, format, or decision boundary.
Fine-tuning
Math & training
Continuing training from pretrained parameters on a narrower dataset or objective. Depending on the method, you may update all parameters, selected parameters, or added adapter parameters.
Flaky Test
AI-native development
A test that can pass and fail across equivalent runs without a relevant change to the code or intended test environment.
FlashAttention
Infrastructure & serving
An exact attention algorithm that tiles the computation to reduce transfers between accelerator memory levels while avoiding materialization of the full attention matrix in high-bandwidth memory.
Function Calling
Agents & tools
A provider or application interface through which a model emits a structured request naming a tool and its arguments. Application code validates the request, performs the operation, and can return the result for another model step.
GAN (Generative Adversarial Network)
Models & inference
A generator network tries to create realistic data while a discriminator network tries to tell real from fake. They train together: the generator gets better at fooling the discriminator, and the discriminator gets better at detecting fakes.
GPT
Models & inference
Generative Pre-trained Transformer, a family label for generative transformer models pretrained on sequence-prediction objectives and adapted for downstream use. Product names and model architectures should not be treated as interchangeable.
Goodput
Infrastructure & serving
The rate of completed requests that satisfy defined service constraints, such as both time-to-first-token and per-token latency objectives, under a stated workload.
Graceful Degradation
Reliability & operations
Preserving a bounded core service when capacity or dependencies are impaired by reducing optional quality, features, freshness, or workload instead of failing every request.
Gradient
Math & training
A vector of partial derivatives pointing in the direction of steepest increase. In ML, you go opposite to the gradient (gradient descent) to minimize the loss.
Gradient Accumulation
Math & training
Summing or averaging gradients from several microbatches before performing one optimizer update.
Gradient Clipping
Math & training
Limiting gradient values or their combined norm before an optimizer update when they exceed a chosen threshold.
Gradient Descent
Math & training
A family of optimization updates that move parameters using the negative gradient of an objective, usually estimated from batches rather than the entire dataset.
Grounding
Retrieval & generation
Connecting a generated answer or action to evidence, state, or observations that the system can identify and check.
Guardrails
Evaluation & safety
System controls that constrain inputs, tool use, outputs, permissions, and escalation. They can include schemas, policy checks, classifiers, allowlists, sandboxing, approvals, and post-action verification.
HNSW
also: Hierarchical Navigable Small World
Retrieval & generation
An approximate-nearest-neighbor index that organizes vectors in layered proximity graphs and searches from coarse upper layers toward detailed lower layers.
Hallucination
Evaluation & safety
Generated content that is false, unsupported by the available evidence, or inconsistent with the task's source of truth. It can arise even when the output is fluent and the model is not attempting to deceive.
Handoff
AI-native development
A structured transfer of a task between people or agents that preserves the objective, current state, evidence, decisions, constraints, and remaining work.
Human-in-the-Loop (HITL)
also: Human oversight, human review
Agents & tools
A workflow design in which a person supplies judgment, correction, approval, or escalation at defined points in an AI-driven process.
Hybrid Retrieval
Retrieval & generation
Retrieval that combines signals from different methods, commonly lexical matching and dense-vector similarity, before merging or reranking results.
Hyperparameter
Math & training
A configuration choice that shapes model structure, optimization, data processing, or inference rather than being learned as an ordinary model parameter. Examples include learning rate, batch size, layer count, and decoding settings.
Idempotency
AI-native development
The property that repeating the same operation with the same identity does not create additional side effects beyond the first successful application.
Image Token
Multimodal systems
A model-specific visual unit represented as a vector or discrete code, commonly derived from an image patch, region, or learned visual-codebook entry.
In-Context Learning
Prompting & context
A model adapting its behavior from instructions, examples, or patterns supplied in the current input without an ordinary parameter update.
Incident Response
Reliability & operations
The coordinated process for detecting, analyzing, containing, recovering from, communicating, and learning from an event that threatens service, data, safety, or security.
Indirect Prompt Injection
Security & governance
A prompt-injection attack delivered through content the system retrieves or observes, such as a webpage, document, email, image text, or tool result, rather than directly through the user's instruction.
Inductive Bias
Models & inference
Structural or statistical assumptions that favor some functions or representations over others. Convolution favors locality and shared filters; causal masking favors prediction from preceding positions.
Inference
Models & inference
Executing a trained model to produce predictions, scores, embeddings, or generated tokens without performing an ordinary training update to its parameters.
Instruction Following
Prompting & context
A model capability to map natural-language directions and supplied context to behavior that satisfies the stated task and constraints.
Instruction Hierarchy
Prompting & context
A rule set for resolving conflicts among instructions from sources with different authority, such as application policy, users, and untrusted retrieved content.
Inter-Token Latency (ITL)
Infrastructure & serving
The elapsed time between two consecutive output-token arrival events for one request, calculated as `t_i - t_(i-1)` for an output token after the first.
JAX
Math & training
A Python library for transforming numerical functions with automatic differentiation, compilation, vectorization, and parallel execution across accelerators. Its transformations work best with explicit state and functional-style code.
Jailbreak
Security & governance
An adversarial input or interaction strategy intended to make a model produce behavior that its training or application controls are designed to prevent.
KV Cache
Models & inference
Stored key and value tensors from earlier positions in autoregressive generation. Reusing them avoids recomputing attention projections for the unchanged prefix at every decoding step.
Knowledge Distillation
Math & training
Training a student model to reproduce selected behavior or output distributions from a more capable teacher, often alongside ordinary target labels.
LLM (Large Language Model)
Models & inference
A language model with enough capacity and broad training to perform many language tasks through prompting or adaptation. Most current LLMs use transformer architectures and sequence-prediction objectives, but size thresholds, data sources, and training recipes vary.
LLM-as-a-Judge
Evaluation & safety
Using a language model to score, compare, classify, or critique another system's output against a rubric.
Late Fusion
Multimodal systems
Processing modalities through separate encoders or predictors and combining their high-level representations, scores, or decisions near the task output.
Latent Space
Data & representations
A learned representation space whose coordinates encode factors useful to a model. It may be lower-dimensional than the input, but compression is not required for every latent representation.
Learning Rate
Math & training
A scale factor used by an optimizer to control parameter-update magnitude. Values that are too large can destabilize training; values that are too small can make useful progress impractically slow.
Learning Rate Schedule
Math & training
A policy that changes the optimizer's learning rate as training progresses according to steps, epochs, metrics, or a predefined curve.
Least Privilege
Evaluation & safety
Giving a model, agent, tool, or user only the permissions required for the current task, for only as long as those permissions are needed.
LoRA (Low-Rank Adaptation)
Math & training
A method that keeps base weights frozen and learns low-rank update matrices for selected layers. It reduces the number of trainable parameters and can lower training memory relative to full-parameter fine-tuning.
Load Shedding
Reliability & operations
Deliberately rejecting, dropping, or cancelling selected work at one or more overload boundaries when demand exceeds the capacity available to produce useful results.
Logits
Models & inference
The model's unnormalized numeric scores for candidate outcomes before a normalization function or decoding rule converts them into selections.
Loss Function
Math & training
An objective that maps predictions and targets, sometimes with regularization terms, to a value optimization tries to reduce. The loss determines which errors training directly rewards or penalizes.
Lost in the Middle
Prompting & context
A long-context failure pattern in which model performance changes with evidence position and can degrade when relevant information sits between the beginning and end.
MCP (Model Context Protocol)
Agents & tools
An open JSON-RPC protocol for a host to connect to servers that expose tools, resources, prompts, and extensions through defined request, result, discovery, and transport contracts. In revision 2026-07-28, every request carries its protocol version and client capabilities instead of relying on an initialization handshake or protocol session.
Maximum Marginal Relevance (MMR)
Retrieval & generation
A selection rule that balances relevance to the query with novelty relative to items already selected.
Membership Inference
Security & governance
An attack that estimates whether a particular record or example was included in a model's training data by observing model outputs or other accessible signals.
Mixed Precision
Math & training
A numerical strategy that uses different data types for different operations, often lower precision for many matrix operations and higher precision for values that need more range or stability.
MoE (Mixture of Experts)
Models & inference
An architecture with multiple expert subnetworks and a learned router that selects a subset for each input unit, often each token. Sparse activation can increase total parameter capacity without using every expert on every forward pass.
Modality
Multimodal systems
A form of information with its own structure and acquisition process, such as text, image, audio, video, depth, or sensor measurements.
Modality Alignment
Multimodal systems
Learning or establishing correspondences between representations from different modalities so semantically or temporally related items can be matched.
Model Card
Evaluation & safety
A structured report describing a model's intended uses, evaluation conditions, performance characteristics, limitations, and relevant ethical or safety considerations.
Model Router
AI-native development
A component that selects a model or provider for a request using requirements such as capability, latency, cost, context size, policy, and current availability.
Model Serving
Infrastructure & serving
The runtime and API layer that loads versioned model artifacts, accepts inference requests, schedules execution, manages resources, and returns results under an operational contract.
Multi Round-Trip Request (MRTR)
also: MRTR
Agents & tools
An MCP request pattern in which an operation returns `resultType: input_required` with one or more `inputRequests`, then the client retries the original method with `inputResponses` and the exact returned `requestState`.
Multimodal Fusion
Multimodal systems
Combining evidence or learned representations from more than one modality to produce a joint representation, prediction, or generated output.
Multimodal Model
Multimodal systems
A model that learns from, relates, or generates more than one modality through representation, alignment, fusion, translation, or coordinated prediction.
NaN (Not a Number)
Math & training
A floating-point value representing an undefined or unrepresentable numerical result. In training, NaNs can come from invalid operations, overflow, unstable normalization, excessive updates, or earlier corrupted values.
Normalization
Math & training
A family of transformations that rescale or recenter inputs, activations, or features using defined statistics. Batch normalization and layer normalization use different axes and behave differently across training and inference.
Nucleus Sampling (Top-p)
also: Top-p sampling
Models & inference
A decoding method that samples from the smallest set of next-token candidates whose cumulative probability reaches a chosen threshold.
Observability
AI-native development
The ability to understand an AI system's behavior from recorded inputs, outputs, state transitions, tool calls, timings, costs, errors, and evaluation signals.
Optimizer
Math & training
An algorithm that transforms gradients into parameter updates. Plain stochastic gradient descent is a simple baseline; momentum, Adam, and other optimizers change the update using history or adaptive scaling. Each choice has different memory, stability, and tuning behavior.
Orchestration
Agents & tools
The control logic that sequences, branches, delegates, retries, pauses, resumes, and terminates work across model and tool steps.
Overfitting
Math & training
A generalization gap in which performance on training data is substantially better than performance on representative unseen data. Memorization can contribute, but the operational symptom is poor generalization.
Paged KV Cache
Infrastructure & serving
A KV-cache memory manager that stores attention state in fixed-size blocks and maps logical sequence positions to physical blocks instead of requiring one contiguous allocation per sequence.
Parameter
Models & inference
A value learned during training, commonly a weight, bias, embedding element, or normalization parameter. Parameter count is one measure of model capacity, but it does not directly determine quality, memory, or serving cost.
Pass@k
Evaluation & safety
Across a task set, the fraction of tasks for which at least one of k sampled candidates passes a defined correctness test.
Patch
AI-native development
A reviewable representation of changes to one or more files, usually expressed as additions and deletions against a known base revision.
Patch Embedding
Multimodal systems
A learned projection that converts an image patch into a fixed-width vector used as one element of a transformer input sequence.
Perplexity
Models & inference
The exponentiated average negative log-likelihood under a stated tokenization and logarithm convention. Lower values mean the model assigned higher probability to the evaluated sequence.
Pipeline Parallelism
Infrastructure & serving
Partitioning sequential groups of model layers across devices and moving microbatches or requests through those stages as a pipeline.
Planning
Agents & tools
Constructing, selecting, or revising a sequence of actions and dependencies intended to move from the current state to a goal.
Postmortem
Reliability & operations
A durable incident record that explains impact, detection, response, contributing conditions, recovery, and owned follow-up actions without assigning blame as a substitute for analysis.
Precision & Recall
Evaluation & safety
Precision asks how many flagged items were correct; recall asks how many relevant items were found. When you change the decision threshold for one fixed scoring model, improving recall often lowers precision and vice versa. A better model can improve both. F1 is their harmonic mean.
Prefill
also: Prefill Phase
Infrastructure & serving
The initial inference stage that processes all supplied input tokens to produce their representations and the attention state required for subsequent autoregressive generation.
Prefix Caching
Infrastructure & serving
Reusing KV-cache blocks produced for an identical eligible token prefix across requests so the serving runtime can skip repeated prefix computation.
Progressive Disclosure
AI-native development
Supplying a person or model with the minimum useful context first, then revealing deeper detail when the task or evidence requires it.
Prompt Cache
Prompting & context
Reuse of provider-side or application-side computation for an identical or eligible prompt prefix so repeated inference avoids some preprocessing work.
Prompt Engineering
Prompting & context
Designing model-facing instructions, examples, constraints, and output requirements to improve behavior on a defined task.
Prompt Injection
Evaluation & safety
An attack or failure mode in which untrusted content influences a model to disregard intended instructions, expose data, misuse tools, or take actions outside the user's goal. The content can arrive directly from a user or indirectly through retrieved pages, files, messages, or tool output.
Prompt Sensitivity
Prompting & context
Variation in model output or measured performance caused by changes to prompt wording, order, formatting, or examples that preserve the intended task.
Provenance Attestation
Security & governance
Authenticated, machine-readable metadata that binds an artifact to claims about how, where, when, and from which inputs it was produced.
Purpose Limitation
Security & governance
For personal data, collecting and using it only for specified, explicit purposes unless a new use has an appropriate compatible or authorized basis.
QLoRA
Math & training
A parameter-efficient fine-tuning method that keeps a pretrained base model frozen in a low-bit quantized representation while training LoRA adapters with higher-precision computation where needed.
Quantization
Models & inference
Representing weights, activations, or caches with lower-precision formats to reduce memory, bandwidth, or compute cost. Methods differ in calibration, granularity, data type, and whether conversion happens before, during, or after training.
RAG (Retrieval-Augmented Generation)
Retrieval & generation
A system pattern that retrieves evidence relevant to a request and supplies selected content to a generative model before it answers or acts. Retrieval can use lexical, vector, structured, or hybrid methods.
RLHF (Reinforcement Learning from Human Feedback)
Math & training
A family of pipelines that uses human feedback to learn a reward or preference signal and then optimizes a model policy against that signal. Implementations vary and need not all use the same reinforcement-learning algorithm.
ROUGE
Evaluation & safety
A family of metrics that compares generated text with reference text using units such as n-gram overlap or longest common subsequence.
Rate Limit
AI-native development
A policy that caps requests, tokens, concurrent work, or another resource within a defined time or capacity window.
ReAct
Agents & tools
An agent pattern that interleaves task reasoning, a concrete action, and an observation returned by the environment before deciding the next step.
ReLU
Math & training
Rectified Linear Unit, defined as `f(x) = max(0, x)`. It is inexpensive and has a non-saturating positive branch, though zero gradients on negative inputs can create inactive units.
Readiness Probe
Reliability & operations
A diagnostic that tells the traffic-routing layer whether a service instance is currently able to accept requests.
Recall@K
Retrieval & generation
For one query, Recall@K is `|relevant items intersecting the top k| / |relevant items|`. A dataset score aggregates those per-query values under a stated rule.
Reciprocal Rank Fusion (RRF)
Retrieval & generation
A rank-fusion method that combines several result lists by summing contributions that decrease with each item's rank in each list.
Red Teaming
Security & governance
A structured adversarial testing process in which authorized testers seek failures using documented objectives, threat assumptions, cases, and evidence.
Regression Test
AI-native development
A repeatable check that protects behavior known to work, especially after code, prompt, model, retrieval, or tool changes.
Repository Instructions
AI-native development
Version-controlled guidance that tells coding agents how a repository is organized, which commands and conventions apply, what boundaries to respect, and how to verify work.
Repository Map
AI-native development
A compact, maintained description of a repository's important directories, ownership boundaries, entry points, build commands, tests, generated files, and local instructions.
Reproducible Build
AI-native development
A build whose declared source, environment, and instructions can be independently rerun to produce bit-for-bit identical specified artifacts.
Reranker
Retrieval & generation
A second-stage model or scoring function that reorders a small candidate set using a richer comparison between the query and each candidate.
Retry Budget
Reliability & operations
A bound on retry traffic, usually expressed relative to original requests or over a time window, that prevents retries from consuming unbounded capacity.
Retry with Backoff
AI-native development
Repeating a failed transient operation after progressively longer delays, usually with randomized jitter and a strict retry limit.
Reviewer Agent
AI-native development
An agent assigned to inspect another agent's artifact or decision against explicit criteria and return findings or a verdict.
Rollback
Reliability & operations
Restoring a previously known deployment or configuration when the current release violates operational, quality, or safety criteria.
SFT (Supervised Fine-Tuning)
Math & training
Fine-tuning a pretrained model on paired inputs and desired responses so it learns the demonstrated behavior under the training distribution.
Sandbox
Agents & tools
An isolated execution environment that restricts an agent's access to files, processes, network destinations, credentials, and host resources.
Saturation
Reliability & operations
The degree to which a constrained resource or service has exhausted its capacity, including queued work that cannot begin promptly.
Scope Contract
AI-native development
A concrete agreement that defines a task's goal, allowed and forbidden surfaces, expected artifacts, verification requirements, and stopping conditions.
Self-Attention
Models & inference
Attention in which queries, keys, and values are derived from the same sequence representation. Scaled similarity scores are normalized and used to combine values, subject to causal, padding, local, or other masks.
Semantic Cache
AI-native development
A cache that reuses a previous result when a new request is judged sufficiently similar under a chosen representation and threshold.
Separation of Duties
Security & governance
Dividing conflicting responsibilities or authority across independent roles so one principal cannot complete a high-risk action without another authorized decision.
Service Level Indicator (SLI)
Reliability & operations
A quantitative measure of service behavior at a defined user-relevant boundary, such as successful request ratio or latency below a threshold.
Service Level Objective (SLO)
Reliability & operations
A target range or threshold for a service-level indicator over a stated population and measurement window.
Shadow Traffic
Reliability & operations
A copy of live request traffic sent to a candidate system for observation while the candidate response remains outside the primary user response path. Because the copied request still executes, its side effects must be isolated.
Shared Embedding Space
Multimodal systems
A common vector space in which representations from different modalities can be compared with the same similarity function.
Skill Bundle
Agents & tools
The complete installable skill directory, including `SKILL.md` and every reference, script, asset, fixture, or companion file required by the workflow.
Skill Catalog
Agents & tools
The compact model-visible inventory of eligible skills, usually containing routing metadata such as name, description, and an internal source identifier rather than every skill body.
Skill Discovery
Agents & tools
A runtime pipeline that searches configured roots, identifies candidate skill directories, validates their package contract, attaches scope and provenance, resolves collisions, and publishes eligible catalog entries.
Skill Invocation
Agents & tools
The runtime-mediated process in which an eligible human, model, application, or other skill selects a skill and causes its instructions to enter the working context.
Softmax
Math & training
A function defined by `softmax(x_i) = exp(x_i) / sum(exp(x_j))`, implemented with numerical stabilization. Its outputs are positive and sum to one, so they can parameterize a categorical distribution.
Software Bill of Materials (SBOM)
also: SBOM
Security & governance
A structured inventory of software components and relationships associated with a product or artifact, often including versions, suppliers, licenses, and identifiers.
Speculative Decoding
Models & inference
An inference method in which a cheaper draft process proposes several tokens and the target model scores those draft positions in parallel. In exact sampling variants, an acceptance and correction rule preserves the target model's output distribution.
Stateless MCP
Agents & tools
The MCP 2026-07-28 request model in which every request carries the protocol version and client capabilities in `params._meta`, while results carry an explicit `resultType`; no protocol state is keyed by an initialization handshake, connection, or `Mcp-Session-Id`.
Stochastic Gradient Descent (SGD)
also: SGD
Math & training
An optimizer family that updates parameters from a gradient estimated on a sampled example or minibatch rather than the complete training dataset.
Stop Sequence
Models & inference
An application-specified token or text pattern that causes generation to stop when the decoding system encounters it.
Streaming
Models & inference
Delivering incremental response events before the complete result is ready. A stream may contain token text, structured deltas, tool-call arguments, usage metadata, or status events depending on the API.
Structured Output
Agents & tools
Model output constrained or validated against a machine-readable schema so application code can consume fields without parsing free-form prose.
Swarm
Agents & tools
A loosely coordinated multi-agent pattern in which local agent decisions and message exchange produce system-level behavior. The term is used inconsistently, so the actual topology, state ownership, and termination rules must be specified.
System Prompt
Prompting & context
A provider-defined instruction message or configuration supplied by the application to establish behavior and constraints within that provider's instruction hierarchy.
Tail Latency
Reliability & operations
The latency experienced by the slowest portion of requests, commonly summarized with a high percentile under a stated workload and time window.
Temperature
Models & inference
A decoding parameter that rescales logits before a probability distribution is formed. Higher positive values usually flatten the distribution; lower positive values sharpen it.
Tensor
Data & representations
A typed array with a shape, data type, and device placement that frameworks use to represent inputs, parameters, activations, and gradients. Automatic-differentiation metadata is framework- and operation-dependent, not an inherent property of every tensor.
Tensor Parallelism
Infrastructure & serving
Partitioning tensor operations within a model layer across devices, with collective communication combining partial results during the layer computation.
Termination Condition
Agents & tools
An explicit rule that ends or pauses an agent run when it succeeds, fails, exhausts a budget, reaches a safe boundary, or requires escalation.
Test Oracle
AI-native development
The mechanism, specification, reference, invariant, or human judgment used to decide whether observed program behavior is correct.
Threat Model
Security & governance
A documented account of protected assets, trust boundaries, potential adversaries, assumed capabilities, attack paths, impacts, and planned controls.
Time per Output Token (TPOT)
Infrastructure & serving
For one request with `N > 1` output tokens, the average post-first-token interval: `(t_N - t_1) / (N - 1)`. System distributions then aggregate those per-request averages.
Time to First Token (TTFT)
also: TTFT
Models & inference
The elapsed time from submitting a generation request until the client receives the first output token or content event under a defined measurement boundary.
Token
Data & representations
An integer identifier produced by a model-specific tokenizer from text, bytes, images, audio, or another input representation. A token can be a whole word, part of a word, punctuation, whitespace, a byte sequence, or a special control symbol.
Token Budget
Prompting & context
An explicit allocation of token capacity across instructions, evidence, history, tool results, reasoning or working space, and output.
Tokenization
Data & representations
Converting an input representation into the ordered token identifiers a specific model or tokenizer accepts.
Tokens per Second (TPS)
also: TPS, output token throughput
Infrastructure & serving
A throughput measure reporting how many output tokens a serving system produces per unit time under a stated scope and workload.
Tool Contract
Agents & tools
The complete agreement for a tool boundary: purpose, typed inputs, outputs, validation, permissions, side effects, errors, timeouts, idempotency, and evidence returned to the caller.
Top-k Sampling
Models & inference
A decoding method that restricts the next-token distribution to the k highest-scoring candidates, renormalizes their probabilities, and samples from that set.
Trace
AI-native development
A correlated record of one request or task across model calls, retrieval, tools, state transitions, retries, approvals, and evaluations.
Transfer Learning
Math & training
Starting from representations or parameters learned on one data distribution or objective and adapting them for another. The transferable components and update strategy depend on architecture and task.
Transformer
Models & inference
A neural-network architecture built from attention, position information, feed-forward sublayers, residual connections, and normalization. Encoder, decoder, and encoder-decoder variants use different masks and information flows.
Trust Boundary
Security & governance
An interface where data, instructions, identity, or authority crosses between components or principals that operate under different trust assumptions.
Underfitting
Math & training
A model or training setup has insufficient effective capacity, optimization, features, or training signal to capture useful patterns in the training data.
VAE (Variational Autoencoder)
Models & inference
A latent-variable model trained with a reconstruction objective and a regularization term that keeps an approximate posterior close to a chosen prior. The reparameterization estimator allows gradients through stochastic latent sampling.
Vector Database
Retrieval & generation
A storage and indexing system that supports nearest-neighbor queries over vector representations, often with metadata filtering, persistence, and approximate indexes.
Verification Gate
Evaluation & safety
A control point that blocks progress until defined evidence satisfies a correctness or quality criterion.
Vision Transformer (ViT)
Multimodal systems
A vision architecture that represents an image as a sequence of patch embeddings with position information and processes that sequence with transformer encoder blocks.
Vision-Language Model (VLM)
Multimodal systems
A model that learns relationships between, or jointly processes, visual and language representations for tasks such as retrieval, description, question answering, or grounded generation.
Visual Grounding
Multimodal systems
Connecting a language expression to spatial evidence in an image or video, such as a region, object, mask, or tracked entity.
Vocabulary
Data & representations
The finite mapping between token identifiers and the units a tokenizer can emit, including ordinary, byte-level, and special control tokens.
Warmup
Math & training
An initial training phase in which the learning rate rises from a smaller value toward the main schedule's target value.
Weight
Math & training
A trainable coefficient in a model transformation. Weights are usually organized into tensors, and optimization adjusts them to reduce the training objective.
Weight Decay
Math & training
An update rule that reduces selected parameter magnitudes over training, often by multiplying weights by a shrinkage factor separate from the gradient update.
Worktree
AI-native development
In Git, a working directory attached to a repository and branch or commit, with shared object storage but its own checked-out files and index.
Zero Trust
Security & governance
A security model that grants no implicit trust from network location or asset ownership and instead evaluates each access request against identity, device, resource, policy, and current context.
Zero-Shot
Prompting & context
Performing a task from instructions or task framing without including task-specific demonstrations in the immediate input.