weastel

Writing / karpenter-midas-touch

Karpenter: Midas Touch for K8s Cluster

The gold you earn may shine, but the gold you save will endure — WarrenGPT

Recently, we cut our EC2 costs by 60% (saving around $160k annually) by upgrading our K8s node provisioner from Cluster Autoscaler (CAS) to Karpenter. This post walks through that move — what we changed, how Karpenter behaves under the hood, and what we measured after migration.

Architecture overview

At Newton School, we run production workloads on Amazon EKS. That cluster backs backend services, analytics and data tooling (Airbyte, Airflow, Metabase), learning infrastructure (on-prem GitLab, Judge Hero, evaluation services), and GPU-heavy AI/ML workloads. In practice that meant 200+ services, 500+ pods, and on the order of 20k users — a mix of spot and on-demand nodes, stateful pieces (Gitaly, databases, caches, queues) and stateless APIs and workers.

We originally used Cluster Autoscaler: node groups plus taints/tolerations to place unscheduled pods. We migrated to Karpenter as a more economical provisioner and saw roughly 60% lower overall cost. The sections below explain how Karpenter works and how we moved without drama.

Karpenter under the hood

Karpenter combines offline and online bin-packing when choosing nodes. Cluster Autoscaler uses a lighter online pack driven mainly by unscheduled pod CPU/memory, bounded by fixed node groups — a smaller menu of instance types.

Bin-packing with cloud constraints is NP-hard, so both tools use approximations. Online packing reacts to new pending pods; offline packing reshapes the fleet by adding or removing nodes to improve utilization.

AWS built Karpenter with price-aware packing: native online and offline optimization with cost as a first-class signal. For a deeper walkthrough, see AWS re:Invent — Harness the power of Karpenter.

Migration strategy

  1. Bootstrap off the cluster — We ran Karpenter on EKS Fargate so Karpenter itself wasn’t scheduled on nodes it manages, with IAM/network permissions to add/remove nodes and update policies.

  2. Node pools — Separate pools for spot vs on-demand, plus GPU pools for ML. For single-replica StatefulSets we disabled aggressive de-provisioning to avoid downtime from offline consolidation.

  3. Cutover — We moved active nodes under Karpenter control with custom policies on existing nodes so the fleet transferred without cluster downtime.

Official guide: Migrating from Cluster Autoscaler.

Interesting insights

Offline bin packing

Without cost-aware offline packing, CAS kept our spot fleet on m5ad.2xlarge when m5.2xlarge would have fit — a >60% spot premium for the same shape of work, driven by node-group rigidity rather than actual pod needs.

Example spot pricing (via vantage.sh at the time of writing):

m5ad.2xlarge spot ≈ $0.2655 / hour
m5.2xlarge spot   ≈ $0.1427 / hour

CPU vs memory assumptions

We started memory-heavy (classic backend/auxiliary services). As products, data, and ML grew, the fleet became more CPU-weighted, but CAS still scaled memory-oriented node groups (m4.2xlarge, m6.2xlarge, etc.). Karpenter’s offline packing let the cluster shift toward CPU-efficient instance choices without us maintaining a perfect node-group matrix.

More scale, more optimization

Spot pricing rises on popular sizes; larger, less contended instances often have better $/core and $/GB. Karpenter isn’t locked to predefined group SKUs, so it can chase better ratios — including big spot nodes when the bin pack says so.

Illustrative math from the original post:

m6a.2xlarge  (8 vCPU, 32 GiB)  spot ≈ $0.1792 / hour
m6a.32xlarge (128 vCPU, 512 GiB) spot ≈ $1.0732 / hour

16 × m6a.2xlarge ≈ 16 × $0.1792 = $2.8672 / hour
1 × m6a.32xlarge ≈ $1.0732 / hour

→ ~⅔ savings for equivalent capacity (when the pack fits)

During CodeRush — a large coding event with 20–30k concurrent users over ~3 hours — Karpenter scaled the fleet for the spike. We saw on the order of ~80% lower hourly cost vs the CAS era for that pattern: roughly $25–30/hour average under CAS vs ~$5.5/hour with Karpenter (charts were in the original Medium post).

Conclusion

Moving from Cluster Autoscaler to Karpenter was a step-change in how we run EKS: cheaper, more flexible node selection, and better behavior at scale. Thanks to the AWS Karpenter team and everyone who made it production-ready in the open.

References

  1. Karpenter
  2. AWS re:Invent 2023 — Harness the power of Karpenter
  3. How Grafana Labs switched to Karpenter (Grafana blog)
  4. Anthropic and Karpenter (The Stack)

Originally published on Medium (Mar 13, 2024).

← All posts