deniz.in

Markets

Weather

Loading weather

· via dev.to (home feed)

Kubernetes 1.37 adds HPA scale-to-zero, gang scheduling beta and drops 18 kubelet flags

Kubernetes 1.37 'Garhwal' lets the HorizontalPodAutoscaler scale workloads to zero replicas, promotes gang scheduling to beta, and removes 18 kubelet flags that can block node startup on upgrade.

Kubernetes 1.37 adds HPA scale-to-zero, gang scheduling beta and drops 18 kubelet flags

Kubernetes 1.37 'Garhwal' lands with 67 enhancements

Kubernetes 1.37 has been available since August 26, 2026, and ships 67 enhancements, 16 of which have graduated to stable, according to dev.to. The release is nicknamed 'Garhwal' after a region in the Indian Himalayas and is framed as a consolidation release: an API that has sat in beta for roughly nine years is finalized, the HorizontalPodAutoscaler gains the ability to shrink workloads to zero replicas, and gang scheduling for distributed AI and ML jobs moves into beta. Operators should also plan around breaking changes that can stop a node from starting after an upgrade.

HPA can now scale workloads to zero

The most operationally significant feature is scale-to-zero in the HorizontalPodAutoscaler, which arrives in beta and is switched on by default. As dev.to explains, the mechanism depends on Object or External metrics, because CPU and memory signals are produced by running pods and vanish once the last replica is gone. A queue depth, by contrast, exists independently of the workers consuming it, so the HPA can keep reading that value even at zero replicas and spin workers back up when work arrives.

The practical payoff is that queue consumers, batch jobs and GPU workloads that sit idle between bursts no longer hold resources while dormant. A new ScaledToZero status on the HPA object makes it clear whether the controller scaled the workload down or a human set the replica count to zero manually.

Gang scheduling and workload-aware preemption

Gang scheduling, tracked as KEP-4671, reaches beta in this release. It introduces the PodGroup concept: the scheduler is told, for example, that eight pods belong to a single training job and must start together, and it waits until enough capacity exists for all of them. This addresses a familiar failure mode in distributed training, where some pods get placed while the rest queue indefinitely, leaving the job stalled while still occupying resources.

A related addition, workload-aware preemption (KEP-5710), stops competing workloads from repeatedly preempting one another in a loop. Together these features are aimed at platform teams running Ray, JobSet or LWS who want a production-grade AI stack on Kubernetes.

metrics.k8s.io finally goes stable

The metrics.k8s.io API has been in beta since Kubernetes 1.8, a wait of nearly nine years, and 1.37 promotes it to stable v1. The API surface itself does not change; this is purely a version graduation, though it signals that the resource metrics infrastructure is considered production-grade. kubectl top now prefers the v1 endpoint and falls back to v1beta1.

Breaking changes: 18 kubelet flags removed

Upgrading to 1.37 requires preparation. The embedded cAdvisor was migrated to a slimmer cadvisor/lib module (PR #139870), which removes 18 kubelet flags, including --containerd, --containerd-namespace, --boot-id-file and --enable-load-reader. A kubelet configured with any of these flags fails to start entirely, reporting an unknown flag, so nodes configured through kubeadm-flags.env or systemd drop-in files must be cleaned up beforehand.

Other removals and requirements noted by dev.to: kube-proxy's legacy IPVS support gives way to nftables, the failCgroupV1 setting stays enabled by default, and container runtimes must run containerd 2.x, since containerd 1.x is no longer supported.

DRA, ulimits and security features mature

Dynamic Resource Allocation receives four stable graduations in this single release, underlining that GPUs, accelerators and special network adapters are now fully integrated into the resource model. A standardized DRA attribute, resource.kubernetes.io/numaNode, enables consistent placement that respects NUMA topology. A per-container ulimits setting via the Container.SecurityContext field arrives in alpha, a feature databases and highly concurrent workloads have long wanted.

On the security side, pod certificates (PodCertificateRequest) are stable: instead of mounting long-lived secrets, the kubelet can request short-lived X.509 certificates and rotate them regularly, effectively giving workloads a native identity system. SELinuxMount is also stable, using the Linux kernel's -o context mount option rather than recursively relabeling an entire volume before each pod starts.

Why it matters

This release shows Kubernetes pulling in two directions at once. On the feature side, scale-to-zero, gang scheduling and matured DRA support all target bursty, GPU-heavy AI workloads, letting idle jobs consume nothing and distributed training run without partial scheduling stalls. On the operational side, the removal of 18 kubelet flags, the IPVS removal and the containerd 2.x requirement mean an unplanned upgrade can leave nodes that refuse to boot. Anyone running a cluster should audit node configuration files before moving to 1.37, while the nine-year journey of metrics.k8s.io to stable finally closes a long-standing gap in the platform's API surface.

  • #kubernetes
  • #containers
  • #autoscaling
  • #cluster-operations
  • #cloud

Related posts