Glossary · Math & training

Weight Decay

An update rule that reduces selected parameter magnitudes over training, often by multiplying weights by a shrinkage factor separate from the gradient update.

Why it matters

It can improve generalization, but the useful coefficient and excluded parameter groups depend on model, optimizer, schedule, and data.

Common confusion

Decoupled weight decay is equivalent to an L2 loss penalty for some simple optimizers, but not generally for adaptive optimizers such as Adam.

Related terms

Browse the learning paths to see this term in context — every lesson is free to read.