Glossary · Evaluation & safety

Guardrails

System controls that constrain inputs, tool use, outputs, permissions, and escalation. They can include schemas, policy checks, classifiers, allowlists, sandboxing, approvals, and post-action verification.

Why it matters

No single filter covers all failure modes, so controls should be layered according to risk.

Common confusion

Guardrails reduce risk; they do not prove that an AI system is safe.

Related terms

Browse the learning paths to see this term in context — every lesson is free to read.