Glossary · Math & training
Weight Decay
An update rule that reduces selected parameter magnitudes over training, often by multiplying weights by a shrinkage factor separate from the gradient update.
Why it matters
It can improve generalization, but the useful coefficient and excluded parameter groups depend on model, optimizer, schedule, and data.
Common confusion
Decoupled weight decay is equivalent to an L2 loss penalty for some simple optimizers, but not generally for adaptive optimizers such as Adam.
Related terms
Browse the learning paths to see this term in context — every lesson is free to read.