Glossary · Infrastructure & serving
Goodput
The rate of completed requests that satisfy defined service constraints, such as both time-to-first-token and per-token latency objectives, under a stated workload.
Why it matters
Raw throughput can rise while users experience more slow requests. Goodput counts only work that meets the service contract.
In practice
Declare the request distribution and latency thresholds, count only compliant completions, report percentiles beside the aggregate rate, and avoid comparing systems under different objectives.
Common confusion
Goodput is not all completed throughput and is not a universal property of a model. It depends on workload and success thresholds.
Related terms
Sources
Browse the learning paths to see this term in context — every lesson is free to read.