Glossary · Infrastructure & serving

Pipeline Parallelism

Partitioning sequential groups of model layers across devices and moving microbatches or requests through those stages as a pipeline.

Why it matters

It lets models exceed one device's memory, but stage imbalance, pipeline bubbles, activation transfers, and failure coordination affect usable performance.

In practice

Balance stage cost, choose a microbatch schedule, measure idle time and interconnect traffic, and keep model and checkpoint partition metadata versioned.

Common confusion

Pipeline parallelism divides layers by depth. Tensor parallelism divides tensor operations within a layer.

Related terms

Sources

Browse the learning paths to see this term in context — every lesson is free to read.