deniz.in

Markets

Weather

Loading weather

· via dev.to (home feed)

Kubelet reservation changes quietly reshape usable node memory across EKS, GKE and AKS

A dev.to analysis finds recent kubelet reservation changes in EKS, GKE and AKS have moved allocatable node memory, with per-node swings of several GiB and real capacity-planning consequences.

Kubelet reservation changes quietly reshape usable node memory across EKS, GKE and AKS

A node with 16 GiB of RAM may only let pods schedule against about 11.9 GiB of it once kubelet reservations are subtracted — and the three major managed Kubernetes platforms have recently changed how those reservations are computed. That is the starting point of an analysis published on dev.to by kalikys, who worked through the EKS nodeadm source code and the current GKE and AKS documentation and calculated allocatable CPU and memory for instance types across all three clouds. The conclusion: much of the capacity advice floating around, including older blog posts and calculators, no longer matches what the platforms actually reserve.

What the kubelet actually reserves

The scheduler does not see a node's raw hardware. The kubelet subtracts system reservations from total capacity to produce an Allocatable figure, and that is what pods compete for. In the managed services these reservations are increasingly driven by maxPods — the per-node pod limit — which means networking configuration choices now have direct memory consequences.

EKS: prefix delegation raises the reservation

According to the dev.to analysis, EKS reserves 11 MiB of memory per possible pod plus a 255 MiB base. With the default VPC CNI, maxPods tracks the instance's ENI limit, which is 17 on a t3.medium, putting the reservation at 442 MiB and leaving roughly 87% of the node's memory usable. Enabling prefix delegation to fit more pods raises maxPods to 110, and the reservation climbs to 1465 MiB, dropping usable memory to 62%. An m5.large falls from 92% to 81%.

The trade reverses on larger machines: an m6i.4xlarge moves from 95.5% to 97.6% usable with prefix delegation, because the per-pod reservation is small relative to total memory. On a 4 GiB node, memory runs out long before the 110th pod could ever launch, so the extra reservation pays for capacity slots the node will never use.

GKE 1.37 hands memory back

GKE's change runs the other direction. Through version 1.36 it reserved a regressive share of memory: 25% of the first 4 GiB, 20% of the next 4, 10% of the following 8, and 6% up to 128 GiB. From 1.37 on Container-Optimized OS the reservation is the lower of that old value and 15 MiB × maxPods + 500 MiB — about 2.1 GiB at 110 pods, regardless of node size.

The dev.to figures show an n2-standard-16 (64 GiB) dropping from 5.5 GiB reserved to 2.1 GiB, and an n2-standard-32 (128 GiB) from 9.3 GiB to 2.1 GiB, freeing 7.2 GiB for pods. The CPU reservation is now also capped at 1 vCPU, with one exception the author flags: the shared-core e2-medium still has 1060m reserved out of its 2 vCPUs, leaving 940m. The practical implication is that node pools sized against the old tiers may be able to shed nodes after a routine upgrade.

AKS: a 250-pod default hits the cap

Since version 1.29, AKS reserves min(20 MB × maxPods + 50 MB, 25% of RAM). Azure CNI Overlay defaults to 250 max pods, which on a 16 GiB D4s v5 pushes the reservation into the 25% cap: 4 GiB reserved, 74% usable. Setting maxPods to 110 at pool creation cuts the reservation to 2.2 GiB and lifts usable memory to 86% — a 1.8 GiB per-node difference from a single flag. The catch is that maxPods cannot be changed on an existing AKS pool; the only route is a new pool and a workload migration.

What to do about it

The author's recommendations are straightforward. Set maxPods to a level the node can realistically run rather than the maximum the CNI permits. Size pod requests against the Allocatable line in kubectl describe node, minus whatever DaemonSets consume — never against the instance spec sheet. And re-check the numbers after control-plane upgrades, since GKE 1.37 demonstrates that the formulas can change underneath a cluster. For those who want to verify against their own setup, the author also publishes a free calculator covering roughly 2,900 instance types across AWS, GCP and Azure, ranking nodes by monthly cost after reservations are applied; the numbers were computed by script and checked against the vendor sources.

Why it matters

Reservations are invisible at purchase time but decisive at scheduling time. Capacity plans and cost models built on instance specs or out-of-date reference tables can be off by gigabytes per node, in either direction — EKS and AKS take memory away through defaults most teams never inspect, while GKE's upgrade quietly gives it back. The AKS case is the sharpest lesson: a networking default chosen at pool creation has a fixed capacity cost, and correcting it later means rebuilding infrastructure. For anyone doing bin-packing density math, right-sizing exercises or upgrade planning, the only number that matters is Allocatable, and it now needs to be re-validated whenever the platform moves.

  • #kubernetes
  • #capacity-planning
  • #eks
  • #gke
  • #aks

Related posts