Glossary · Infrastructure & serving

Inter-Token Latency (ITL)

The elapsed time between two consecutive output-token arrival events for one request, calculated as `t_i - t_(i-1)` for an output token after the first.

Why it matters

Individual gaps expose decode stalls and streaming jitter that a per-request average can hide, especially under batching, preemption, or mixed workloads.

In practice

Record each post-first-token interval with its request and token position, then report distributions by workload, output length, and concurrency without pooling away request boundaries.

Common confusion

ITL is one interval between consecutive tokens. Time per output token is a per-request average across those intervals, while time to first token covers the wait before streaming begins.

Related terms

Sources

Browse the learning paths to see this term in context — every lesson is free to read.