Kubernetes Autoscaling

Autoscaling on Kubernetes happens at two layers. Pod autoscaling changes how many replicas a workload runs, or how big each one is. Node autoscaling changes how many machines the cluster has so those Pods have somewhere to run. The two work together: the Horizontal Pod Autoscaler adds Pods, some of them can't be scheduled, and the node autoscaler adds a node for them.

Done well, autoscaling keeps latency flat during spikes and cuts cloud costs during quiet hours. Done badly, it flaps, lags behind traffic, or scales on a metric that doesn't reflect load at all.

TL;DR

Quick Example

Scale an API between 3 and 30 replicas to hold average CPU near 70% of requests, with calmer scale-down:

Scale-up can double the replica count every 30 seconds. Scale-down waits five minutes of sustained low load, then removes at most two Pods a minute.

Core Concepts

Horizontal Pod Autoscaler

The HPA controller runs a loop (every 15 seconds by default) that computes:

Metric types:

Custom and external metrics need an adapter such as the Prometheus Adapter or a cloud metrics adapter, or KEDA, which serves the external metrics API for you. With several metrics configured, the HPA scales to whichever yields the most replicas.

Vertical Pod Autoscaler

The VPA watches actual usage and recommends requests. In Off mode it only publishes recommendations (the safe way to start). In Initial mode it sets requests at Pod creation. In Recreate/Auto mode it evicts Pods to apply new sizes; with in-place Pod resize (now GA), some changes can apply without a restart. VPA is great for right-sizing batch jobs and singletons, and for finding sensible requests for everything else.

KEDA

KEDA (Kubernetes Event-Driven Autoscaling) adds a ScaledObject resource with dozens of scalers: Kafka consumer lag, RabbitMQ, SQS, Azure Service Bus, Redis lists, Prometheus queries, cron windows, and more. It drives an HPA under the hood and adds scale-to-zero: with no events, the workload runs zero replicas, and KEDA scales it back up when messages arrive. It's the natural fit for message queue consumers and background jobs.

Node Autoscaling

Choosing What to Scale On

Best Practices

Get Requests Right First

HPA utilization is a percentage of requests. Requests set far too high means the HPA never scales; far too low means it scales constantly. Run VPA in recommendation mode or review real usage before tuning HPA targets.

Leave Headroom and Tune Behavior

Target 60–75% CPU, not 95%. New Pods take time to start and warm up, and node provisioning can take a minute or more. Scale up aggressively and scale down slowly with a stabilization window to avoid flapping.

Keep a Sensible Minimum

minReplicas should cover baseline traffic and survive losing a zone. Scale-to-zero suits async workers, not latency-sensitive APIs, because the first request pays for a cold start.

Budget Disruptions

Node consolidation drains nodes. PodDisruptionBudgets keep enough replicas running while it happens, and karpenter.sh/do-not-disrupt or safe-to-evict annotations protect long-running jobs.

Common Mistakes

HPA and VPA Fighting Over CPU

If you use both, have the HPA scale on a custom metric such as RPS, and let the VPA manage only resources the HPA doesn't use.

No Resource Requests at All

Without CPU requests, the HPA can't compute utilization and reports <unknown>. Nothing scales.

Scaling Pods Without Scaling Nodes

An HPA that creates Pods the cluster can't schedule just produces a pile of Pending Pods. Pair Pod autoscaling with Cluster Autoscaler or Karpenter, and set maxReplicas within what your node limits and budget allow.

FAQ

How fast does the HPA react?

The controller evaluates every 15 seconds, metrics-server scrapes about every 15 seconds, and new Pods then need to start and pass readiness. Expect 30–90 seconds for Pods to add capacity, plus node provisioning time if new nodes are needed. For predictable spikes, pre-scale on a schedule (KEDA's cron scaler works well).

Should I use Cluster Autoscaler or Karpenter?

On AWS, Karpenter is the common choice for new clusters: it provisions faster, picks cheaper instance types, and consolidates automatically. Cluster Autoscaler is mature, works on every major cloud, and fits well if you already manage node groups carefully. Managed options like GKE Autopilot and EKS Auto Mode handle node scaling for you.

Can Kubernetes scale to zero?

Not with the built-in HPA, whose minimum is one replica. KEDA and Knative both scale to zero and back up on demand. It's ideal for queue workers and low-traffic internal services.

Why isn't my HPA scaling?

Run kubectl describe hpa. Common causes: metrics-server isn't installed, Pods lack resource requests, the custom metrics adapter isn't serving the metric, or the workload is already at maxReplicas.

Related Topics

References