Glossary · Infrastructure & serving

Disaggregated Serving

A serving architecture that runs prefill and decode work in separately provisioned worker pools and transfers the required attention state between them.

Why it matters

Prefill and decode stress hardware differently, so independent pools can be sized and scheduled for their own bottlenecks instead of competing in one queue.

In practice

Measure state-transfer cost, route requests through compatible model versions, scale each pool from its own demand signal, and test failure recovery between phases.

Common confusion

Disaggregation separates runtime stages. It does not split one model into tensor or pipeline-parallel shards within a stage.

Related terms

Sources

Browse the learning paths to see this term in context — every lesson is free to read.