· via dev.to (home feed)
Bridge add-on versions to roll EKS nodes across Kubernetes upgrades without downtime
A dev.to post shows how moving EKS add-ons to versions valid on both the old and new Kubernetes releases closes the risky gap in AWS's recommended upgrade order.

A post on dev.to lays out a technique for upgrading Amazon EKS node groups without downtime: before rolling nodes from one Kubernetes minor to the next, move every EKS add-on to the newest version that is supported on both the current and the target Kubernetes release. The author calls this the "bridge version" trick, and it addresses a specific weakness in the upgrade sequence AWS documents.
The gap in the standard upgrade order
According to the dev.to post, the AWS upgrade guide prescribes the order control plane, then nodes, then add-ons. That works on paper, but it leaves a window during the node roll where new nodes are running old add-on versions that were only validated against the previous Kubernetes minor.
The add-ons that hurt most are the ones every pod depends on. Vpc-cni hands out pod IPs on every node, CoreDNS resolves names for every workload, and aws-ebs-csi-driver attaches volumes for stateful pods. If one of them misbehaves on a freshly rolled node, the author notes, you find out in the middle of the rollout with pods rescheduling all around you.
What a bridge version is
A bridge version is the newest add-on version that appears in the supported-version lists for both your current and your target Kubernetes release. If every add-on is on its bridge version before the nodes roll, old nodes run an add-on version validated for the old release, new nodes run the same version validated for the new release, and no node ever runs an add-on that was not validated for its Kubernetes version.
Because a bridge version is valid on both releases, the post says it can be applied at any point before the node upgrade: at the very start while still on the older minor, or right after the control plane moves up. The order the author uses is control plane, then bridge add-ons, then nodes, then optionally settling on the default add-on versions. This does not contradict AWS's guidance; it inserts one extra safe step ahead of the node roll.
Finding the bridge version
EKS can list every available version of an add-on for a given Kubernetes version through aws eks describe-addon-versions. The bridge is simply the newest version present in both lists, and the post includes a bash script that queries both, intersects them, and picks the highest common entry, warning that an intermediate hop may be needed if no overlap exists.
On the author's cluster, moving from Kubernetes 1.35 to 1.36, most add-ons already had an identical latest version on both releases: vpc-cni at v1.23.1-eksbuild.1, coredns at v1.14.6-eksbuild.4, aws-ebs-csi-driver at v1.66.0-eksbuild.1 and metrics-server at v0.9.0-eksbuild.11. For those, the bridge is just the latest build.
Kube-proxy is the exception. Its version tracks the Kubernetes minor, so the 1.36 latest (v1.36.0-eksbuild.25) is not available on 1.35, and the bridge is the 1.35 build v1.35.3-eksbuild.29. The advice is to keep kube-proxy on that bridge until the nodes are fully on 1.36.
How the Terraform is wired
The control plane version comes from a single variable, var.cluster_version, so changing it from 1.35 to 1.36 makes Terraform update the cluster in place. The post stresses that EKS upgrades one minor version at a time and offers no rollback, making this the step to plan most carefully. The VPC and access configuration does not change during an upgrade, and nodes and add-ons are separate resources, so the control plane can be targeted on its own.
The managed node group reads the same version variable, which ties the node AMI to the control plane version. Changing the variable rolls the group onto the matching EKS-optimized AMI. Two settings shape how that roll behaves: max_unavailable = 1 rolls one node at a time so the rest keep serving traffic, and force_update_version = true lets the update proceed even when a PodDisruptionBudget blocks pod drains. The author flags this as the one setting that can cost availability, since EKS will eventually evict pods despite the PDB, so critical workloads need enough replicas.
Add-ons are managed with a single aws_eks_addon resource iterating over a list variable, so bumping versions is a data change. The EBS CSI driver is the only add-on that receives a service account role ARN.
Why it matters
The node roll is where an EKS upgrade becomes real: that is when workloads actually move and when a broken add-on surfaces as an incident. Shifting add-on compatibility work to a bridge version, applied before the roll, removes the version mismatch that makes node upgrades risky, and it slots into the AWS-recommended order rather than replacing it. For teams managing EKS with Terraform, the approach also keeps the upgrade path simple: one variable drives the control plane and nodes, and add-on versions become data changes with known-good targets. The caveats remain the same as ever: kube-proxy will not have a same-latest version on both releases, forced updates can override disruption budgets, and the control plane bump cannot be undone.
- #aws-eks
- #kubernetes
- #terraform
- #cloud
- #devops