Glossary · Math & training

Stochastic Gradient Descent (SGD)

Also known as: SGD

An optimizer family that updates parameters from a gradient estimated on a sampled example or minibatch rather than the complete training dataset.

Why it matters

It is the baseline for understanding gradient noise, momentum, batch scaling, and the adaptive optimizers used in modern training.

In practice

Record batch sampling, learning rate, momentum if used, and schedule, then compare validation behavior under equal update or token budgets.

Common confusion

In current practice, SGD usually means minibatch SGD, and its useful learning rate does not follow one universal batch-scaling rule.

Related terms

Sources

Browse the learning paths to see this term in context — every lesson is free to read.