Projects

kOps to EKS Migration

The Kubernetes control plane was ours to runrun by AWS.

Led by my team lead and a colleague. I rehearsed the IP move, built the QA cluster and rewrote the server automation.

About 270 single-tenant customer servers ran on self-managed kOps clusters in five AWS regions.

Why it was needed

We owned the control plane, etcd and every upgrade, and each upgrade meant downtime for single-replica servers. The hard part was the network: customers allow our inbound and outbound IP addresses in their firewalls, so new addresses would have meant a firewall change at every customer.

AWS region · one of fiveinboundCustomer firewallsallow our two IPsInbound Elastic IPcarried acrossAmazon EKScontrol plane run by AWS~270 serversrestored by VeleroNAT gatewaysame outbound IP
Both Elastic IPs were carried across from the kOps clusters to EKS. Customer data came across by Velero backup and restore.

How it was done

  1. Moved the inbound Elastic IP between Terraform states with state rm and import, with a published plan of one change and nothing destroyed.

  2. Copied every server with Velero, scaled to zero first for a clean snapshot, and kept the source volumes for rollback.

  3. Ran all customers in a region at once with one Python script, one region per window, smallest first.

Problems on the way

  • AWS cannot detach an Elastic IP from a NAT gateway.

    Freeing the outbound address meant deleting the gateway, which cuts outbound traffic. Fix: Saved the Terraform plan before the delete, so recovery was one reviewed apply.

  • The smallest region exposed an availability-zone mismatch.

    A carried-over address brings where the old cluster lived, not just its IP. Fix: Stopped and fixed it before moving the larger regions.

Keep the addresses, replace the clusters.

Results

  • ~270customer servers moved, minutes of downtime each
  • 5regions migrated, one per maintenance window
  • 80 minfull dry run of 208 tenants, by one script