Glossary · Models & inference

Inference

Executing a trained model to produce predictions, scores, embeddings, or generated tokens without performing an ordinary training update to its parameters.

Common confusion

An application can update caches, conversation state, or external memory during inference even though model weights stay unchanged.

Related terms

Browse the learning paths to see this term in context — every lesson is free to read.