Glossary · Infrastructure & serving

Tensor Parallelism

Partitioning tensor operations within a model layer across devices, with collective communication combining partial results during the layer computation.

Why it matters

It lets one layer use memory and compute from several devices, but frequent communication can dominate when the interconnect or partition is unsuitable.

In practice

Match partition dimensions to model shapes, benchmark collective traffic, keep ranks on a fast interconnect, and record the sharding layout with checkpoints and serving configuration.

Common confusion

Tensor parallelism splits work inside layers. Pipeline parallelism places different layer groups on different devices.

Related terms

Sources

Browse the learning paths to see this term in context — every lesson is free to read.