The elapsed time from submitting a generation request until the client receives the first output token or content event under a defined measurement boundary.
Why it matters
TTFT strongly affects perceived responsiveness and can reveal queueing, prompt processing, cache, or network delays.
In practice
Record client-side TTFT by model, prompt length, region, and cache status, then separate it from total completion time.
Common confusion
TTFT is not tokens per second. One measures startup latency; the other measures generation throughput after output begins.
Related terms
Browse the learning paths to see this term in context — every lesson is free to read.