Glossary · Infrastructure & serving
Continuous Batching
A serving scheduler that adds and removes generation requests at iteration boundaries instead of waiting for every request in a fixed batch to finish.
Why it matters
Autoregressive requests produce different output lengths, so continuous batching can keep accelerators utilized without forcing short requests to wait for the longest one.
In practice
Admit new requests when capacity becomes available, track per-request latency, and apply backpressure when the live batch or KV-cache budget is full.
Common confusion
Continuous batching is an inference scheduling policy, not gradient accumulation or a training batch-size technique.
Related terms
Sources
Browse the learning paths to see this term in context — every lesson is free to read.