Glossary · Infrastructure & serving
Tokens per Second (TPS)
Also known as: TPS, output token throughput
A throughput measure reporting how many output tokens a serving system produces per unit time under a stated scope and workload.
Why it matters
It complements startup latency by showing how quickly generation proceeds after output begins and how serving behaves under load.
In practice
State whether TPS is per request or aggregate, exclude or identify prefill, and report batch, concurrency, sequence lengths, hardware, and percentile latency.
Common confusion
TPS is not directly comparable across different tokenizers, workloads, quality settings, or measurement boundaries.
Related terms
Sources
Browse the learning paths to see this term in context — every lesson is free to read.