Glossary · Infrastructure & serving
Model Serving
The runtime and API layer that loads versioned model artifacts, accepts inference requests, schedules execution, manages resources, and returns results under an operational contract.
Why it matters
A capable model can still produce an unreliable product when queueing, batching, placement, versioning, cancellation, and response boundaries are not engineered explicitly.
In practice
Pin model and tokenizer versions, validate request limits, expose readiness and latency signals, control concurrency, and test rollback before routing production traffic.
Common confusion
Model serving is broader than calling inference once and narrower than the complete application, which may also include retrieval, tools, policy, and user state.
Related terms
Sources
Browse the learning paths to see this term in context — every lesson is free to read.