Glossary · Math & training

Gradient Accumulation

Summing or averaging gradients from several microbatches before performing one optimizer update.

Why it matters

It lets you approximate a larger effective batch when one device cannot hold all examples and activations at once.

In practice

Scale the loss consistently, call the optimizer only after the chosen number of microbatches, and measure whether normalization or distributed synchronization changes behavior.

Common confusion

Gradient accumulation reduces per-step activation memory, but it does not reproduce every property of processing the full batch simultaneously.

Related terms

Sources

Browse the learning paths to see this term in context — every lesson is free to read.