Glossary · Models & inference
Inference
Executing a trained model to produce predictions, scores, embeddings, or generated tokens without performing an ordinary training update to its parameters.
Common confusion
An application can update caches, conversation state, or external memory during inference even though model weights stay unchanged.
Related terms
Browse the learning paths to see this term in context — every lesson is free to read.