deniz.in

Markets

Weather

Loading weather

· via dev.to (home feed)

Deleting AWS Copilot stacks can delete databases, DNS and load balancers

AWS Copilot is deprecated, and the standard rebuild-in-Terraform-then-delete migration can destroy databases, certificates and DNS records unless resources are protected first.

Deleting AWS Copilot stacks can delete databases, DNS and load balancers

Copilot is deprecated, but its stacks live on

AWS dropped support for the Copilot CLI on June 12, 2026 and archived the project's repository ten days later, according to a detailed write-up on dev.to. Services deployed with Copilot keep running, because the tool was never a runtime: it generated CloudFormation stacks, and those stacks outlived the CLI.

That leaves teams running ECS workloads with no supported management tool. The migration most reach for is to rebuild everything in Terraform, then remove the old Copilot stacks. Done naively, the author warns, that final step takes production down with it, because parts of Copilot's scaffolding are wired to destroy shared resources during deletion.

Four deletion traps

After reading Copilot's source code, the author identified four mechanisms that make teardown dangerous:

  1. Lambda-backed custom resources. Copilot ships 13 of them, and 9 run destructive code when Delete fires. Removing an environment stack triggers handlers that remove the ACM certificate together with its DNS validation records, strip every Route 53 alias record for custom domains, delete the NS delegation record from the app's hosted zone, and purge the ELB access-logs bucket, including all object versions. Importing those resources into Terraform beforehand does not save you: the Lambda runs regardless of which tool now manages the resource.

  2. The env-controller. Every service stack contains an EnvControllerAction custom resource. Once the last service that needs a shared ALB, NAT gateways or EFS is deleted, the env-controller rewrites the environment stack to remove them, so the load balancer and file system disappear alongside the service.

  3. Unprotected addons. Aurora, DynamoDB and S3 addons live in a nested stack with no DeletionPolicy, so deleting the parent stack lets CloudFormation take the database with it.

  4. The --retain-resources flag does not help in the way people assume. Per the write-up, it only functions on stacks already stuck in DELETE_FAILED.

Retain, import, then delete

The safe sequence, per the author, reverses the usual order. First, apply DeletionPolicy: Retain and UpdateReplacePolicy: Retain to every resource in every stack, nested stacks and the StackSet included; retaining a Custom::* resource also stops CloudFormation from invoking its Delete handler. Second, import each resource into Terraform with its exact deployed values, so the first plan is import-only. Only then delete the Copilot stacks: CloudFormation simply forgets the resources and nothing is destroyed.

ecsodus, an alpha tool for the migration

Applying those policies by hand across dozens of resources invites mistakes, so the author built ecsodus, an Apache-2.0 CLI that automates exactly this sequence. Design choices the author highlights: every AWS client is wrapped in a guard that permits only Describe, List, Get and Lookup calls, so the tool cannot modify anything; the Terraform it generates is built from literal deployed values, and a check phase rejects any plan containing creates, updates, deletes or replacements; the retain patch edits templates via change sets that are accepted only if they change deletion policies and nothing else; and no stack is ever split, so anything unsupported stays on Copilot with an explanation in the report. Reports also lead with the zero-risk baseline of simply keeping the CloudFormation as-is.

On September 30 the author ran the tool end to end in a sandbox account against a real Copilot v1.34.1 app: an environment, a Load Balanced Web Service, and DynamoDB and S3 addons, five stacks and 49 resources in total. The reported outcome was 44 of 44 targeted resources imported with zero changes and zero destruction, all Copilot stacks plus the StackSet deleted, the service still answering with HTTP 200, sentinel data in DynamoDB and S3 intact, and the env-controller and rule-priority Lambdas never invoked during teardown.

The tool is explicitly alpha. Version 0.1 covers Load Balanced Web Services, Backend Services, their environment, and Aurora/RDS, DynamoDB and S3 addons. Worker Services, Scheduled Jobs, Request-Driven Web Services, Static Sites, NLB, CloudFront, sidecars and pipelines are detected and reported as blocked, with App Runner planned next. The author is asking remaining Copilot users to run the read-only inventory and report commands and open issues describing what is blocked, to shape the roadmap.

Why it matters

Deprecating a CLI does not decommission the infrastructure it created, and Copilot's cleanup logic was designed for teardown, not handoff. That makes the seemingly routine "delete the old stacks" step of a rebuild-and-cutover migration the exact step that can erase production databases, certificates and DNS. The retain-first, import-first sequence reframes this as an adoption in place rather than a rebuild, and tooling that enforces it by construction matters when a single missed DeletionPolicy has irreversible consequences. One caveat: the results come from the tool author's own sandbox test of alpha software, so treat them as a starting point and run the read-only inventory against your own stacks before touching anything.

  • #aws
  • #cloudformation
  • #terraform
  • #aws-copilot
  • #migration

Related posts