Glossary · Infrastructure & serving

Model Serving

The runtime and API layer that loads versioned model artifacts, accepts inference requests, schedules execution, manages resources, and returns results under an operational contract.

Why it matters

A capable model can still produce an unreliable product when queueing, batching, placement, versioning, cancellation, and response boundaries are not engineered explicitly.

In practice

Pin model and tokenizer versions, validate request limits, expose readiness and latency signals, control concurrency, and test rollback before routing production traffic.

Common confusion

Model serving is broader than calling inference once and narrower than the complete application, which may also include retrieval, tools, policy, and user state.

Related terms

Sources

Browse the learning paths to see this term in context — every lesson is free to read.