Glossary · Infrastructure & serving
Pipeline Parallelism
Partitioning sequential groups of model layers across devices and moving microbatches or requests through those stages as a pipeline.
Why it matters
It lets models exceed one device's memory, but stage imbalance, pipeline bubbles, activation transfers, and failure coordination affect usable performance.
In practice
Balance stage cost, choose a microbatch schedule, measure idle time and interconnect traffic, and keep model and checkpoint partition metadata versioned.
Common confusion
Pipeline parallelism divides layers by depth. Tensor parallelism divides tensor operations within a layer.
Related terms
Sources
Browse the learning paths to see this term in context — every lesson is free to read.