Glossary · Security & governance
Jailbreak
An adversarial input or interaction strategy intended to make a model produce behavior that its training or application controls are designed to prevent.
Why it matters
Successful jailbreaks expose gaps between stated policy and actual behavior, and they can become more consequential when the model controls tools or protected data.
In practice
Derive test families from prohibited behaviors, vary format and interaction length, measure both refusal and harmful completion, and convert confirmed failures into versioned adversarial evals.
Common confusion
A jailbreak targets model or system behavioral restrictions. Prompt injection redirects instruction following, often toward an attacker's goal; one interaction can involve both.
Related terms
Sources
Browse the learning paths to see this term in context — every lesson is free to read.