kOps to EKS Migration
The Kubernetes control plane was ours to runrun by AWS.
Led by my team lead and a colleague. I rehearsed the IP move, built the QA cluster and rewrote the server automation.
About 270 single-tenant customer servers ran on self-managed kOps clusters in five AWS regions.
Why it was needed
We owned the control plane, etcd and every upgrade, and each upgrade meant downtime for single-replica servers. The hard part was the network: customers allow our inbound and outbound IP addresses in their firewalls, so new addresses would have meant a firewall change at every customer.
How it was done
Moved the inbound Elastic IP between Terraform states with state rm and import, with a published plan of one change and nothing destroyed.
Copied every server with Velero, scaled to zero first for a clean snapshot, and kept the source volumes for rollback.
Ran all customers in a region at once with one Python script, one region per window, smallest first.
Problems on the way
AWS cannot detach an Elastic IP from a NAT gateway.
Freeing the outbound address meant deleting the gateway, which cuts outbound traffic. Fix: Saved the Terraform plan before the delete, so recovery was one reviewed apply.
The smallest region exposed an availability-zone mismatch.
A carried-over address brings where the old cluster lived, not just its IP. Fix: Stopped and fixed it before moving the larger regions.
Keep the addresses, replace the clusters.
Results
- ~270customer servers moved, minutes of downtime each
- 5regions migrated, one per maintenance window
- 80 minfull dry run of 208 tenants, by one script