Glossary · Math & training
AdamW
An Adam variant that decouples weight decay from the gradient-based parameter update. That makes the shrinkage behavior easier to reason about than adding an L2 penalty inside Adam's adaptively scaled gradient.
Common confusion
Decoupled weight decay does not make AdamW universally optimal. Model, data, and training scale still determine the best optimizer and schedule.
Related terms
Sources
Browse the learning paths to see this term in context — every lesson is free to read.