Glossary · AI-native development
Rate Limit
A policy that caps requests, tokens, concurrent work, or another resource within a defined time or capacity window.
Why it matters
It protects providers and your own system from overload, uncontrolled spend, and unfair resource use.
In practice
Enforce per-tenant token and concurrency limits, read provider retry metadata, and queue or reject excess work predictably.
Common confusion
A rate limit controls allowed usage. Backpressure propagates downstream capacity constraints through a system.
Related terms
Browse the learning paths to see this term in context — every lesson is free to read.