Glossary · Infrastructure & serving

Goodput

The rate of completed requests that satisfy defined service constraints, such as both time-to-first-token and per-token latency objectives, under a stated workload.

Why it matters

Raw throughput can rise while users experience more slow requests. Goodput counts only work that meets the service contract.

In practice

Declare the request distribution and latency thresholds, count only compliant completions, report percentiles beside the aggregate rate, and avoid comparing systems under different objectives.

Common confusion

Goodput is not all completed throughput and is not a universal property of a model. It depends on workload and success thresholds.

Related terms

Sources

Browse the learning paths to see this term in context — every lesson is free to read.