Glossary · Infrastructure & serving

Continuous Batching

A serving scheduler that adds and removes generation requests at iteration boundaries instead of waiting for every request in a fixed batch to finish.

Why it matters

Autoregressive requests produce different output lengths, so continuous batching can keep accelerators utilized without forcing short requests to wait for the longest one.

In practice

Admit new requests when capacity becomes available, track per-request latency, and apply backpressure when the live batch or KV-cache budget is full.

Common confusion

Continuous batching is an inference scheduling policy, not gradient accumulation or a training batch-size technique.

Related terms

Sources

Browse the learning paths to see this term in context — every lesson is free to read.