Glossary · Evaluation & safety

Prompt Injection

An attack or failure mode in which untrusted content influences a model to disregard intended instructions, expose data, misuse tools, or take actions outside the user's goal. The content can arrive directly from a user or indirectly through retrieved pages, files, messages, or tool output.

Why it matters

Models process instructions and data through the same language channel, so input filtering alone cannot reliably separate every malicious instruction from legitimate content.

In practice

Treat external content as untrusted, isolate it from authority-bearing instructions, minimize tool permissions, require approval for consequential writes, and verify outputs and actions.

Common confusion

Prompt injection is not technically the same mechanism as SQL injection, and a stronger system prompt is not a complete defense.

Related terms

Sources

Browse the learning paths to see this term in context — every lesson is free to read.