← Blog

Kubernetes Cost Optimization: 10 Tactics That Actually Work

Real-world strategies our DevOps team has used to cut Kubernetes infrastructure costs by 40%+ without sacrificing reliability or performance.

Priya Reddy · 2025-01-08 · Cloud

Kubernetes Cost Optimization: 10 Tactics That Actually Work

Kubernetes infrastructure costs spiral because teams provision for peak load and never revisit their allocations. After running cost optimization engagements for 20+ clients, here are the 10 tactics with the highest ROI — plus how to measure savings, a 30-day plan for working through them, and the mistakes that quietly undo the wins.

Why does my Kubernetes bill keep growing?

Kubernetes makes it easy to request capacity and very quiet about waste. A developer copies a deployment manifest with 2 CPU / 4Gi requests from another service, the scheduler dutifully reserves that capacity on a node, and the pod actually uses 200 millicores. Multiply by a few hundred pods and the cluster runs at 15–25% real utilisation while the cloud bill reflects 100% of the reserved nodes. Three structural facts drive this: requests (not usage) determine node count; nobody owns the gap between the two; and dev/staging environments replicate production sizing without production traffic.

The Big 3 — which Kubernetes cost fixes come first?

  1. Right-size your requests and limits — 70% of clusters we see are over-provisioned by 2x or more. Use Goldilocks or the Vertical Pod Autoscaler in recommendation mode to compare each workload's requests against its actual P95 usage over a couple of representative weeks, then cut requests to P95 plus sensible headroom. This is pure waste removal: no architecture change, no user-visible risk if you move in steps.
  2. Enable cluster autoscaling — match node capacity to actual load instead of running for peak 24/7. Most teams we audit keep 30–40% idle nodes around the clock for traffic that arrives two hours a day. Autoscaling only works when step 1 is done: over-stated requests make the autoscaler think the cluster is full when it isn't.
  3. Move stateless workloads to Spot/Preemptible instances — 60–80% cheaper for interruption-tolerant work (API replicas behind a load balancer, queue consumers, CI runners, batch jobs). Handle the interruptions properly: PodDisruptionBudgets, more than one instance type in the node pool, and graceful shutdown hooks so a reclaimed node drains instead of dropping requests.
deployment.yaml — what right-sizing actually looks like
# Before: copied from another service two years ago, never revisited
resources:
  requests: { cpu: "2", memory: 4Gi }
  limits:   { cpu: "2", memory: 4Gi }

# After: P95 usage over two weeks was 210m CPU / 900Mi memory
resources:
  requests: { cpu: 300m, memory: 1200Mi }
  limits:   { memory: 1200Mi }   # memory limit only — no CPU limit, no throttling

The Next 7 — what else cuts a Kubernetes bill?

  1. Namespace-level resource quotas — cap dev/staging so an experiment can't quietly triple the cluster. Quotas turn cost conversations from archaeology into a pull request.
  2. Delete idle environments — preview and dev namespaces nobody has touched for 7+ days. A TTL controller or a scheduled job that flags (then removes) stale namespaces routinely claws back 10–15% of the bill.
  3. Use Karpenter instead of the classic Cluster Autoscaler on AWS — it provisions right-sized nodes directly from pending pods in seconds, picks cheaper instance types automatically, and consolidates underused nodes without manual node-group tuning.
  4. Adopt ARM nodes (Graviton and equivalents) — typically ~20% cheaper for comparable performance. Most mainstream runtimes and base images ship multi-arch today; start with stateless services whose images already publish arm64 variants.
  5. Optimise image size — multi-stage builds and slim base images cut pull times (faster scale-up, which lets autoscaling run tighter), registry storage, and cross-zone egress. Going from a 1.2GB image to 150MB is common and free.
  6. Buy reserved capacity or savings plans for the baseline — after right-sizing, a steady floor of usage remains; committing to it for 1 year typically saves 30–40% versus on-demand. Do this last, not first: committing to today's inflated baseline locks the waste in.
  7. Cost allocation tagging and showback — label workloads by team/product and put a weekly cost report in front of the owners. You can't optimise what you can't attribute, and in practice visibility alone changes engineer behaviour within a sprint or two.

How do you measure Kubernetes cost savings?

Pick your metrics before touching anything, or you won't be able to prove the savings. The two that matter: cost per namespace/team (from OpenCost or Kubecost, both of which map cloud billing onto Kubernetes objects) and cluster utilisation — actual CPU/memory usage divided by allocatable capacity. A healthy production cluster after optimisation typically runs 50–65% utilisation; below 30% means you're paying for air. Snapshot both for two weeks before the first change.

First look at the request-vs-usage gap, no tooling required
# Actual usage right now, biggest consumers last
kubectl top pods -n payments --no-headers | sort -k2 -h | tail

# What those same pods reserved from the scheduler
kubectl get pods -n payments -o custom-columns=\
'NAME:.metadata.name,\
CPU_REQ:.spec.containers[*].resources.requests.cpu,\
MEM_REQ:.spec.containers[*].resources.requests.memory'

# Cluster-level: allocatable vs requested per node
kubectl describe nodes | grep -A5 'Allocated resources'

What does a realistic 30-day plan look like?

  1. Week 1 — visibility: deploy OpenCost/Kubecost, label workloads by owner, snapshot utilisation and per-team cost. No changes yet.
  2. Week 2 — right-size the top 20 workloads by reserved capacity, using two weeks of usage data. This alone usually cuts 20–30%.
  3. Week 3 — turn on autoscaling (or migrate to Karpenter), add PodDisruptionBudgets, and move the first batch of stateless workloads to Spot.
  4. Week 4 — quotas on non-production, idle-environment cleanup, and only now price reserved capacity for the remaining steady baseline.

What undoes Kubernetes cost savings?

  • Buying reserved instances before right-sizing — the single most common way to lock waste in for a year.
  • Setting CPU limits far below requests 'for safety' — causes throttling incidents that get answered by re-inflating requests everywhere.
  • Running Spot without disruption handling — one bad reclaim event and the team swears off Spot forever, forfeiting the biggest single discount available.
  • Treating optimisation as a one-off project — costs drift back within a quarter without quotas, showback, and someone owning the utilisation number.

Related Reading