Glossary · Math & training

AdamW

An Adam variant that decouples weight decay from the gradient-based parameter update. That makes the shrinkage behavior easier to reason about than adding an L2 penalty inside Adam's adaptively scaled gradient.

Common confusion

Decoupled weight decay does not make AdamW universally optimal. Model, data, and training scale still determine the best optimizer and schedule.

Related terms

Sources

Browse the learning paths to see this term in context — every lesson is free to read.