Glossary · Infrastructure & serving
Disaggregated Serving
A serving architecture that runs prefill and decode work in separately provisioned worker pools and transfers the required attention state between them.
Why it matters
Prefill and decode stress hardware differently, so independent pools can be sized and scheduled for their own bottlenecks instead of competing in one queue.
In practice
Measure state-transfer cost, route requests through compatible model versions, scale each pool from its own demand signal, and test failure recovery between phases.
Common confusion
Disaggregation separates runtime stages. It does not split one model into tensor or pipeline-parallel shards within a stage.
Related terms
Sources
Browse the learning paths to see this term in context — every lesson is free to read.