Glossary · Reliability & operations

Tail Latency

The latency experienced by the slowest portion of requests, commonly summarized with a high percentile under a stated workload and time window.

Why it matters

Averages can look healthy while a meaningful group of users waits much longer because of queueing, contention, retries, or variable request cost.

In practice

Report several percentiles by route and workload, retain timeouts as censored or failed observations according to a documented rule, and trace slow requests across dependencies.

Common confusion

Tail latency is not the single slowest request and has no meaning without the percentile, population, and measurement boundary.

Related terms

Sources

Browse the learning paths to see this term in context — every lesson is free to read.