Glossary · Math & training
Gradient Accumulation
Summing or averaging gradients from several microbatches before performing one optimizer update.
Why it matters
It lets you approximate a larger effective batch when one device cannot hold all examples and activations at once.
In practice
Scale the loss consistently, call the optimizer only after the chosen number of microbatches, and measure whether normalization or distributed synchronization changes behavior.
Common confusion
Gradient accumulation reduces per-step activation memory, but it does not reproduce every property of processing the full batch simultaneously.
Related terms
Sources
Browse the learning paths to see this term in context — every lesson is free to read.