Glossary · Infrastructure & serving

Autoscaling

A control loop that changes the number or capacity of serving workers from observed demand, resource use, or application metrics within configured bounds.

Why it matters

AI workloads can change faster than manual provisioning, but scaling decisions must account for model-load time, accelerator availability, queueing, and request cost.

In practice

Scale from a demand signal tied to useful work, set minimum warm capacity, bound scale-down churn, and verify that new replicas pass readiness checks before receiving traffic.

Common confusion

Autoscaling adds or removes capacity. It does not make an overloaded dependency faster or guarantee that enough hardware can be acquired in time.

Related terms

Sources

Browse the learning paths to see this term in context — every lesson is free to read.