Glossary · Infrastructure & serving

Expert Parallelism

Distributing mixture-of-experts subnetworks across devices and routing each token's activations to the devices that host its selected experts.

Why it matters

Sparse experts increase model capacity without executing every expert for every token, but routing introduces communication, load-balance, and placement constraints.

In practice

Measure token distribution by expert, provision communication bandwidth, cap or route overflow deliberately, and test quality when traffic produces uneven expert demand.

Common confusion

Expert parallelism partitions experts selected by a router. Tensor parallelism partitions the tensor operations inside layers.

Related terms

Sources

Browse the learning paths to see this term in context — every lesson is free to read.