Glossary · AI-native development

Semantic Cache

A cache that reuses a previous result when a new request is judged sufficiently similar under a chosen representation and threshold.

Why it matters

It can reduce latency and cost for repeated intents, but an incorrect match can return stale or user-inappropriate output.

In practice

Cache low-risk FAQ answers by normalized intent, include tenant and policy version in the key, and bypass the cache for personalized or time-sensitive requests.

Common confusion

Semantic similarity does not guarantee that two requests have the same correct answer. A semantic cache reuses a prior result, while prefix caching reuses exact-token KV state and prompt caching follows provider or application eligibility rules.

Related terms

Browse the learning paths to see this term in context — every lesson is free to read.